Electronic device, method, and non-transitory computer-readable storage medium for generating stereoscopic image or stereoscopic video by using alpha channel including depth value

By adding depth values to an alpha channel and controlling display assemblies, the electronic device enhances stereoscopic image and video rendering, addressing the limitations of existing AR devices and providing improved immersion.

WO2025244282A1PCT designated stage Publication Date: 2025-11-27SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/004635
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-09-19
Filing Date
2025-04-04
Publication Date
2025-11-27

AI Technical Summary

Technical Problem

Existing augmented reality devices struggle to provide immersive stereoscopic experiences by effectively incorporating depth information into images and videos, limiting user engagement and realism.

Method used

An electronic device that adds depth values to an alpha channel of an image, allowing for the generation of stereoscopic images and videos by controlling display assemblies to shift displayed portions based on these depth values, thereby enhancing depth perception.

Benefits of technology

The solution provides an immersive user experience by accurately rendering three-dimensional objects, improving user engagement and realism in augmented reality environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025004635_27112025_PF_FP_ABST
    Figure KR2025004635_27112025_PF_FP_ABST
Patent Text Reader

Abstract

An electronic device comprises: a memory including one or more storage media for storing instructions; and at least one processor including a processing circuit, wherein the instructions, when executed individually or collectively by the at least one processor, cause the electronic device to: acquire depth information of a visual object; identify depth values on the basis of the depth information of the visual object; add the depth values to an alpha channel of an image representing the visual object; and generate the image, and the alpha channel includes the depth values, and transparencies of the visual object.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device, method, and non-transitory computer-readable storage medium for generating a stereoscopic image or stereoscopic video using an alpha channel containing depth values

[0001] The present disclosure relates to an electronic device, a method, and a non-transitory computer-readable storage medium for generating stereoscopic images and / or stereoscopic videos using an alpha channel containing depth values.

[0002] To provide an enhanced user experience, electronic devices are being developed to provide augmented reality (AR) services that display computer-generated information in conjunction with external objects in the real world. These electronic devices may be wearable devices worn by users, such as AR glasses or head-mounted devices (HMDs).

[0003] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above is applicable as prior art related to the present disclosure.

[0004] According to one aspect of the present disclosure, an electronic device includes a memory including one or more storage media storing instructions, and at least one processor including a processing circuit, wherein the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to obtain depth information of a visual object, identify depth values ​​based on the depth information of the visual object, add the depth values ​​to an alpha channel of an image representing the visual object, and generate the image, wherein the alpha channel includes the depth values ​​and transparencies of the visual object.

[0005] According to one aspect of the present disclosure, a method of an electronic device includes an operation of obtaining depth information of a visual object, an operation of identifying depth values ​​using the depth information, an operation of adding the depth values ​​to an alpha channel of an image representing the visual object, and an operation of generating the image, wherein the alpha channel includes the depth values ​​and transparencies of the visual object.

[0006] According to an embodiment, an electronic device may include a memory storing instructions and including one or more storage media, and at least one processor including a processing circuit. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain depth information of a visual object. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify, using the depth information, depth values ​​to be included in an alpha channel representing transparency of the visual object. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate an image representing the visual object, wherein the depth values ​​and transparencies are each included in an alpha channel of pixels.

[0007] In one embodiment, a method of an electronic device may be provided. The method may include an operation of obtaining depth information of a visual object. The method may include an operation of using the depth information to identify depth values ​​to be included in an alpha channel representing transparency of the visual object. The method may include an operation of generating an image representing the visual object, wherein the depth values ​​and transparency values ​​are each included in an alpha channel of pixels.

[0008] In one embodiment, a non-transitory computer-readable storage medium storing instructions may be provided. The instructions, when executed by an electronic device including a display assembly including a plurality of displays, may cause the electronic device to obtain a file including an image representing a visual object. The instructions, when executed by the electronic device, may cause the electronic device to identify, from an alpha channel of pixels of the image, depth values ​​and transparencies of portions of the visual object corresponding to each of the pixels. The instructions, when executed by the electronic device, may cause the electronic device to control the plurality of displays such that, while controlling the display assembly to display a visual object represented based on the transparencies, portions displayed on a first display of the plurality of displays are shifted from portions displayed on a second display of the plurality of displays, respectively, according to the depth values.

[0009] In one embodiment, an electronic device may include a display assembly including a plurality of displays, a memory storing instructions and including one or more storage media, and at least one processor including a processing circuit. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain a file including an image representing a visual object. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify, from an alpha channel of pixels of the image, depth values ​​and transparencies of portions of the visual object corresponding to each of the pixels. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to control the plurality of displays such that, while controlling the display assembly, the portions displayed on a first display of the plurality of displays are shifted from the portions displayed on a second display of the plurality of displays, respectively, according to the depth values, to display the visual object expressed based on the transparencies.

[0010] The above-described and other aspects, features, and advantages of some embodiments of the present disclosure will become more apparent from the following description taken in conjunction with the accompanying drawings, in which:

[0011] FIGS. 1A and 1B illustrate exemplary operations of an electronic device performing three-dimensional rendering using an alpha channel of two-dimensional pixels, according to one embodiment;

[0012] FIG. 2 illustrates a block diagram of an electronic device according to one embodiment;

[0013] FIGS. 3A and 3B illustrate programs running on an electronic device, according to one embodiment;

[0014] FIGS. 4A and 4B illustrate a flow diagram of an electronic device according to one embodiment;

[0015] FIG. 5 illustrates an exemplary operation of an electronic device performing scaling for depth values;

[0016] FIG. 6 illustrates an exemplary operation of an electronic device for generating a video based on key frames;

[0017] FIG. 7 illustrates an exemplary operation of an electronic device that generates a video representing the motion of a visual object using sensor data;

[0018] FIG. 8 illustrates an exemplary operation of an electronic device performing three-dimensional rendering of a visual object represented by an image file;

[0019] FIG. 9A illustrates an example of a perspective view of an electronic device, according to one embodiment;

[0020] FIG. 9B illustrates an example of one or more hardwares disposed within an electronic device, according to one embodiment;

[0021] FIGS. 10A and 10B illustrate an example of an appearance of an electronic device according to one embodiment;

[0022] Fig. 11 shows an example of a block diagram of an electronic device; and

[0023] Fig. 12 shows an example of a block diagram of an electronic device for displaying an image in a virtual space.

[0024] Hereinafter, one or more embodiments of the present disclosure are described with reference to the accompanying drawings.

[0025] FIGS. 1A and 1B illustrate exemplary operations of an electronic device (101) performing three-dimensional rendering using an alpha channel of two-dimensional pixels, according to one embodiment. The electronic device (101) may include a head-mounted display (HMD) wearable on a user's (105) head. The electronic device (101) may be referred to as a head-mounted display (HMD) device, a headgear electronic device, a glasses-type (or goggle-type) electronic device, a video see-through (VST) device, an extended reality (XR) device, a virtual reality (VR) device, and / or an augmented reality (AR) device.

[0026] FIG. 1A illustrates an external appearance of an electronic device (101) having a form of glasses, but embodiments of the present disclosure are not limited thereto. For example, the electronic device (101) may include a mobile phone (e.g., a smartphone having a bar shape, a foldable phone having a flexible display including a bendable portion), a laptop PC (personal computer), a desktop PC, and / or a tablet PC. An example of a hardware configuration included in the electronic device (101) having various form factors described above is exemplarily described with reference to FIG. 2. FIG. 9A, 9B, FIG. 10A, or FIG. 10B illustrate an example of a structure of an electronic device (101) wearable on the head of a user (105). Since the electronic device (101) is wearable on the head of a user (105), the electronic device (101) may be referred to as a wearable device. The electronic device (101) may include an accessory (e.g., a strap) for attaching to the head of a user (105).

[0027] Referring to FIG. 1A, the electronic device (101) can display a virtual object (140) 'stereoscopically'. That is, the virtual object (140) can be displayed in three dimensions, or in 2.5 dimensions (or pseudo-3 dimensions). Throughout this disclosure, the term "stereoscopically" refers to three dimensions or '2.5' dimensions (pseudo-3 dimensions).

[0028] The virtual object (140) may be described as a graphical object defined using a point cloud, vertices, and / or mesh. In the present disclosure, the virtual object (140) may be referred to as a visual object, a visual element, and / or a virtual element. According to one embodiment, the electronic device (101) may display a stereoscopic image and / or a stereoscopic video representing the virtual object (140) to a user (105) (e.g., a user (105) wearing the electronic device (101)). For example, the electronic device (101) may display an image of the virtual object (140) as seen from a virtual camera spaced apart from the virtual object (140) within a virtual space including the virtual object (140) on the display.

[0029] FIG. 1A illustrates a state of an electronic device (101) displaying an exemplary virtual object (140), referred to as an avatar, an augmented reality (AR) emoticon, a virtual reality (VR) emoticon, an AR emoji, and / or a VR emoji. The avatar can be created to represent a user (105) of the electronic device (101) (or a user associated with the avatar). The avatar can be customized by the user (105) of the electronic device (101). Using the avatar representing the user (105), the electronic device (101) can execute functions related to an online service (e.g., a metaverse, a social network service (SNS), and / or a service based on a digital twin). For example, the electronic device (101) may register an avatar in the online service that expresses the user's (105) reaction (e.g., the user's (105) facial expression and / or emotional reaction) to content (e.g., news, posts, articles, and / or (text) messages) provided through the online service.

[0030] In one embodiment, the electronic device (101) may support a selfie function based on a virtual object (140). For example, the electronic device (101) may detect motions of the user (105) while the electronic device (101) is worn by the user (105) (e.g., motions of the user's (105) head, hands, and / or eyes, or motions of the face, referred to as facial expressions). The electronic device (101) may use the detected motions to change the shape and / or position of the virtual object (140), thereby providing a user experience in which the virtual object (140) mimics the motions of the user (105). The electronic device (101) may support a function of capturing a virtual object (140) that reflects the motions of the user (105). The capturing may be performed based on a virtual camera defined within a virtual space including the virtual object (140). For example, the electronic device (101) may generate or store an image and / or video representing a virtual object (140) having a shape and / or location based on the motion of the user (105). The image and / or video may be stored in a file (110).

[0031] Referring to FIG. 1A, a file (110) may include metadata (120) and pixel data (130). Various information describing the pixel data (130) and / or the file (110) may be stored within the metadata (120), for example, based on a format such as EXIF ​​(exchangeable image file format). The pixel data (130) may include information about pixels of an image and / or video included in the file (110). When generating a file (110) representing an image and / or video representing a virtual object (140), the electronic device (101) may generate pixel data (130) representing the image and / or video, and metadata (120).

[0032] For example, the electronic device (101) can generate raw data based on two-dimensional pixels representing a two-dimensional projection of a virtual object (140). The raw data can include a color, transparency (or opacity or alpha value), and a depth value of each of the pixels. For example, when the color is expressed based on three primary colors of red, green, and blue, the electronic device (101) can obtain five attributes (e.g., brightness (or luminance, intensity, power of each of the three primary colors representing the color), transparency, and a depth value) for each of the pixels. For example, raw data representing an image with a width of w and a height of h can include five values ​​of w Х h Х. When each of the above values ​​is expressed as a binary number of 8 bits (or 1 byte), the electronic device (101) can generate raw data having a size of w Х h Х 5 Х 8 bits.

[0033] According to one embodiment, the electronic device (101) may generate or obtain pixel data (130) from raw data representing a virtual object (140) based on pixels having brightness, transparency, and depth values ​​of each of three primary colors. The pixel data (130) may be set to have four attributes (or elements, or channels) for each of the pixels. The four attributes (or channels) may include a red attribute (or red channel) representing the brightness of red light included in the color of the pixel, a blue attribute (or blue channel) representing the brightness of blue light included in the color of the pixel, a green attribute (or green channel) representing the brightness of green light included in the color of the pixel, and / or an alpha attribute (or alpha channel) representing the transparency of the pixel.

[0034] When each of the values ​​of the above attributes is expressed as an 8-bit binary number, the values ​​may be included in an integer range of 0 to 255. When each of the above values ​​is expressed as an 8-bit binary number, the electronic device (101) may represent the attributes of one pixel (e.g., color, transparency, and depth value) using 32 bits (= 8 bits Х 4). For example, pixel data (130) representing an image having a width of w and a height of h may have a size of w Х h Х 32 bits. The embodiment is not limited thereto. For example, the electronic device (101) may generate pixel data (130) having a size less than w Х h Х 32 bits by applying a compression algorithm (or encoding algorithm). Pixel data (130) to which the compression algorithm is applied may have a size less than w Х h Х 32 bits.

[0035] According to one embodiment, the electronic device (101) may generate pixel data (130) in which only four attributes (e.g., brightness of each of the three primary colors and transparency) are set to be assigned to a single pixel for compatibility. By combining and / or encoding a depth value with a designated attribute among the four attributes (e.g., an attribute to which transparency is set to be assigned), the electronic device (101) may generate pixel data (130) that further includes the depth value of the pixel while maintaining compatibility.

[0036] FIG. 1A illustrates values ​​corresponding to pixels p1 and p2, which are included (or compressed, or decoded) in pixel data (130). For example, the electronic device (101) may insert or add a set of values ​​(r1, g1, b1, α1 + d1) representing pixel p1 into the pixel data (130). For example, the electronic device (101) may record or embed a set of values ​​(r2, g2, b2, α2 + d2) representing pixel p2 from the pixel data (130). Here, '+' may represent a concatenation of bits. In the present disclosure, the concatenation (or concatenation operation) of the first value and the second value may mean an operation that performs a bitwise operation, such as a shift operation, to output a third value (e.g., a concatenated value) in which the first value and the second value are serially connected. For example, the first value 1011 (2) and the second value 0101 (2) The concatenation of the first and second values ​​sequentially includes the third value 10110101 from the MSB (most significant bit). (2) It may mean an operation that outputs. The number of bits of the third value may correspond to the sum of the number of bits of the first value and the second value. The electronic device (101) may perform division and / or parsing on the third value to obtain or identify the first value and the second value from the third value.

[0037] For example, the electronic device (101) may generate or obtain pixel data (130) that includes a set (r1, g1, b1, α1 + d1) including concatenated values ​​of transparency (α1) and depth value (d1) having a size of 8 bits, together with 8-bit values ​​(r1, g1, b1), as information (or vector) corresponding to the pixel p1. FIGS. 3A, 3B, 4A, 5, 6, and / or 7 illustrate exemplary operations of the electronic device (101) to generate pixel data (130) and a file (110) including the pixel data (130), according to one embodiment.

[0038] As described above, in one embodiment, the electronic device (101) generates pixel data (130) and a file (110) including the pixel data (130), although embodiments of the present disclosure are not limited thereto. For example, the electronic device (101) may display a virtual object (140) from the file (110). For example, the electronic device (101) may obtain color, transparency, and / or depth values ​​of each of the pixels by decompressing (or decoding) the pixel data (130).

[0039] For example, an electronic device (101) that identifies four values ​​(r1, g1, b1, α1 + d1) for pixel p1 from pixel data (130) can display pixel p1 such that pixel p1 having a color of (r1, g1, b1) is recognized as having a depth (or distance) d1 from a user (105) wearing the electronic device (101). For example, the electronic device (101) can adjust positions at which pixel p1 is visible to each of the two eyes of the user (105) based on binocular parallax to create a depth perception (e.g., a depth perception corresponding to depth d1) of the user (105) for pixel p1. Similarly, an electronic device (101) that has identified four values ​​(r2, g2, b2, α2 + d2) for pixel p2 can display pixel p2 having a color of (r2, g2, b2) at a location having a binocular disparity corresponding to a depth d2. FIG. 4B and / or FIG. 8 illustrate exemplary operations of an electronic device (101) for displaying and / or visualizing a virtual object (140) from a file (110).

[0040] FIG. 1B illustrates an exemplary state of an electronic device (101) that executes a selfie function. The electronic device (101) worn on the head of a user (105) may include displays (151, 152) arranged to face the two eyes of the user (105). On the displays (151, 152), the electronic device (101) may display a stereoscopic image (165) of the user (105) wearing the electronic device (101). The pixels of the image (165) may have positional differences (e.g., positional differences related to binocular disparity) in each of the displays (151, 152). Based on the positional differences, the user (105) wearing the electronic device (101) may recognize that the image (165) represents his / her face stereoscopically.

[0041] The electronic device (101) can display a virtual object (160) (e.g., a virtual object referred to as a virtual camera and / or viewpoint) for controlling the direction of the face of the user (105) expressed through an image (165) on the displays (151, 152). The electronic device (101) can receive an input for moving the virtual object (160) based on a hand gesture of the user (105), a gaze direction (or information indicating a gaze direction) of the user (105), a touch input on the electronic device (101) (or a remote controller connected to the electronic device (101), and / or a voice input based on speech. The electronic device (101) that receives the input can change the position of the virtual object (160) within the displays (151, 152). The electronic device (101) that receives the above input can at least partially change the image (165) based on the changed position of the virtual object (160). For example, the electronic device (101) can display an image (165) that simulates the face of the user (105) as seen from the virtual position represented by the virtual object (160).

[0042] Within the state of FIG. 1B, which displays an image (165), the electronic device (101) may receive an input (e.g., a photographing input) for capturing the image (165). The electronic device (101) that has received the input may generate or store a file (110) including pixel data (130) and metadata (120). The pixel data (130) may include information about colors of pixels included in the image (165). The pixel data (130) may further include information (e.g., a depth value) for displaying the image (165) in three dimensions. FIG. 1B illustrates a set of values ​​(rm, gm, bm, am + dm) representing a pixel Pm of the pixel data (130). The above set can identify or extract the brightness values ​​(rm, gm, bm) of the three primary colors included in the color of the pixel Pm, and the values ​​(am + dm) in which the transparency and depth values ​​of the pixel Pm are encoded. For example, five types of information (e.g., brightness values, transparency, and depth values ​​of each of the three primary colors) can be encoded in the four values ​​included in the above set.

[0043] As described above, according to one embodiment, the electronic device (101) can generate pixel data (130) and file (110) that support three-dimensional rendering of a virtual object (140) while being readable by other electronic devices that extract red, green, blue, and transparency by adding a depth value to transparency among red, green, blue, and transparency. The electronic device (101) can generate or store pixel data (130) that includes both transparency and depth values ​​based on concatenation. By using the pixel data (130) to three-dimensionally render the virtual object (140), the electronic device (101) can provide an immersive user experience to a user (105) wearing the electronic device (101).

[0044] FIG. 2 illustrates a block diagram of an electronic device (101) according to one embodiment. Referring to FIG. 2, the electronic device (101) may be one of various forms of electronic devices, such as smartphones having various form factors (e.g., a bar-type smartphone (101-1), foldable-type smartphones (101-2, 101-3), or sliderable (or rollable) type smartphones), a tablet PC (personal computer) (101-5), a head-mounted display (HMD) device (101-4), a digital camera (101-6), a watch, a cellular phone, a laptop PC, a desktop PC, and / or other similar computing devices.

[0045] In one embodiment, the electronic device (101) may be referred to as a mobile device, a user equipment (UE) (or user terminal), a multi-function device, a portable communication device, a portable device, or a server. The form factor of the electronic device (101) is not limited to the exemplary form factors illustrated in FIG. 2. For example, the electronic device (101) may be included as an electronic control unit (ECU) in a vehicle (e.g., an electric vehicle (EV)). For example, the electronic device (101) may have a form suitable for displaying images and / or videos.

[0046] Referring to FIG. 2, according to one embodiment, an electronic device (101) may include a processor (210) and / or a memory (220). The electronic device (101) may further include a display (230) and / or a sensor (240). The processor (210) may be electrically and / or operatively coupled with the memory (220) and / or the display (230). Electrical coupling of electronic components may include a state in which a wired signal path (or a connection for wireless communication) for transmitting signals is established between the electronic components. Operational coupling of electronic components may include a state in which the electronic components are directly coupled (or a state in which the electronic components are indirectly coupled) such that one of the electronic components controls another electronic component.

[0047] Referring to FIG. 2, a processor (210) of an electronic device (101) may include circuits (e.g., processing circuits and / or cores) for performing operations (e.g., arithmetic operations and / or logical operations) on data. Binary codes (e.g., instructions) representing the operations may be input to the processor (210). The processor (210) may include a central processing unit (CPU), a graphic processing unit (GPU), and / or a neural processing unit (NPU). The processor (210) may be referred to as an application processor (AP) and / or a system on a chip (SoC). The processor (210) may have a structure (e.g., a multi-core structure based on a combination of multiple core circuits such as a dual core, a quad core, a hexa core, or an octa core) for loading (or fetching) and / or executing multiple instructions simultaneously. Within an electronic device (101) comprising at least one processor, including a processor (210), the at least one processor may individually or collectively perform the operations of the present disclosure. For example, the at least one processor may individually and / or collectively perform the operations of FIG. 4A and / or FIG. 4B by executing instructions stored in the memory (220).

[0048] The memory (220) of FIG. 2 may include a circuit for storing data (or instructions) input to or output from the processor (210). The memory (220) may include volatile memory, such as random-access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM). The non-volatile memory may be referred to as storage. The volatile memory may include, for example, at least one of dynamic RAM (DRAM), static RAM (SRAM), cache RAM, and pseudo SRAM (PSRAM). The non-volatile memory may include, for example, at least one of programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), flash memory, hard disk, compact disc, solid state drive (SSD), and embedded multimedia card (eMMC). The memory (220) may include one or more storage media (e.g., the volatile memory and / or non-volatile memory described above) located in a distributed manner in the electronic device (101). The processor (210) of the electronic device (101) may execute instructions of the memory (220) within the electronic device (101) to perform functions and / or operations indicated by the instructions (e.g., the operations of FIG. 4A and / or FIG. 4B).

[0049] The display (230) of the electronic device (101) may include a circuit for visualizing information provided from the processor (210). The display (230) may include a liquid crystal display (LCD), a plasma display panel (PDP), and / or light emitting diodes (LEDs). The LEDs may include organic LEDs (OLEDs). For example, the display (230) may include electronic paper. For example, if the electronic device (101) includes a lens for transmitting external light (or ambient light), the display (230) may include a projector (or projection assembly) for projecting light onto the lens. The display (230) may be referred to as a display panel and / or a display module. The number of displays (230) included in the electronic device (101) may vary depending on the embodiment. For example, an electronic device (101) in the form of an HMD device (101-4) may include displays positioned over each of the user's two eyes when the HMD device (101-4) is worn by the user (e.g., the user (105) of FIG. 1A and / or FIG. 1B). The combination of the displays included in the HMD device (101-4) may be referred to as a display assembly.

[0050] In one embodiment, the display area (or active area) of the display (230) may include a light-emitting area formed by pixels (e.g., activated pixels) of the display (230). The display (230) may include a sensor (e.g., a touch sensor) for detecting an external object (e.g., a user's finger) on the display (230). The sensor may be included in the display (230) in the form of a panel (e.g., a touch sensor panel (TSP)).

[0051] In one embodiment, the sensor (240) of the electronic device (101) may generate electrical information that may be processed by the processor (210) and / or the memory (220) from non-electronic information related to the electronic device (101). For example, the sensor (240) may include a global positioning system (GPS) sensor for detecting the geographic location of the electronic device (101). In addition to the GPS method, the sensor (240) may generate information indicating the geographic location of the electronic device (101) based on a global navigation satellite system (GNSS) such as, for example, Galileo or Beidou (compass). The information may be stored in the memory (220), processed by the processor (210), and / or transmitted to another electronic device distinct from the electronic device (101) via a communication circuit. In one embodiment, the sensor (240) of the electronic device (101) may include an image sensor for acquiring images and / or videos. In one embodiment, the electronic device (101) may have the form of an HMD device (101-4), and the electronic device (101) may include a plurality of image sensors configured to acquire images of the two eyes, facial expressions, hand gestures, and / or the external environment of a user wearing the HMD device (101-4).

[0052] According to one embodiment, the electronic device (101) may generate an image and / or a video representing an avatar (e.g., a virtual object (140) of FIG. 1A) corresponding to a user using information acquired from a sensor (240) as described above with reference to FIG. 1A and / or FIG. 1B. The image and / or the video may be stored in a file (e.g., file (110) of FIG. 1A and / or FIG. 1B). Within the file, the image and / or the video may be stored based on a format set so that four numerical values ​​(e.g., binary numbers and / or binary codes) are assigned to one pixel. The numerical values ​​may correspond to four channels, each representing an attribute of the pixel. According to one embodiment, the processor (210) of the electronic device (101) may generate or load a file set so as to represent all of the transparency and depth values ​​using an alpha channel among the channels.

[0053] FIG. 3a and / or FIG. 3b illustrate exemplary operations of an electronic device (101) for generating or loading a file configured to represent both transparency and depth values ​​using an alpha channel.

[0054] FIGS. 3A and 3B illustrate programs executed on an electronic device according to one embodiment. The electronic device (101) of FIGS. 1A, 1B, and / or 2 may include the electronic device of FIGS. 3A and 3B. The programs illustrated in FIGS. 3A and 3B may be executed by the electronic device (101) of FIG. 2 and / or the processor (210).

[0055] Referring to FIG. 3A, an electronic device may render a virtual object (e.g., the virtual object (140) of FIG. 1A), and obtain an image (e.g., a two-dimensional image) based on a two-dimensional rendering of the virtual object and / or depth information corresponding to the image (e.g., operation (310)). For example, the depth information may represent depth values ​​of each pixel of the image. The depth information may be determined using a reference value (preliminarily) stored in the electronic device. The reference value may be determined (empirically) using an appropriate depth value to stereoscopically display an image to be played back through a file (110). The electronic device may obtain depth information by changing the reference value according to information detected using a sensor of the electronic device (e.g., information indicating the position, posture, and / or shape of the hand and / or face of a user wearing the electronic device).

[0056] Within operation (312), the electronic device may determine or calculate the transparencies (e.g., alpha values) of the pixels of the image of operation (310). The transparencies may be determined using a reference transparency (pre-stored in the electronic device). For example, the reference transparency may be minimum at the center of the image and increase as the distance from the center of the image increases. In other words, the reference transparency may represent an image in which the center area is opaque and the surrounding area is transparent. Within operation (314), the electronic device may use the depth information of operation (310) to encode depth values ​​represented by the depth information into an alpha channel (A) representing the transparency of the virtual object. For example, the electronic device may use the depth information of operation (310) to identify depth values ​​to be included in the alpha channel (A) representing the transparency of the virtual object. The electronic device may generate an image representing the virtual object in which the depth values ​​and transparencies are each included in the alpha channel (A) of the pixels. The electronic device may generate or store a file (110) containing pixel data representing the image (e.g., pixel data (130) of FIG. 1A).

[0057] In one embodiment, the file header and / or metadata (e.g., metadata (120) of FIGS. 1A and / or 1B) of the file (110) may include information (e.g., a flag value) indicating that an alpha channel (A) includes a depth value. The information may indicate the length (e.g., number of bits) and / or position of the depth value within the alpha channel (A). To maintain compatibility, a default value for recognizing transparency through the alpha channel (A) may be set in the information.

[0058] Referring to FIG. 3A, among the 8 bits included in the alpha channel (A), bits corresponding to transparency and depth information, respectively, are exemplarily described. The electronic device can identify a range of depth values ​​represented by the depth information of the operation (310). Based on the range, the electronic device can determine the number of bits to be occupied for representing depth values ​​and the number of bits to be occupied for representing transparencies within the alpha channel (A). The electronic device can determine a ratio of the number of bits for representing transparency and depth values, respectively, based on characteristics of transparency (e.g., range and / or importance) and / or importance of the depth value. The ratio can be increased or decreased depending on the importance of the transparency and depth value. Based on the determination, the electronic device can generate an image including depth values ​​and transparencies within the alpha channel (A), and / or a file (110) representing the image.

[0059] For example, based on determining that depth values ​​are represented by a first number of bits (e.g., 4), the electronic device can generate an image using the depth information, including depth values ​​represented by the first number of bits, and transparencies represented by a second number of bits (e.g., 4 = 8 - 4) obtained by subtracting the first number of bits from the total number of bits included in the alpha channel (A) (e.g., 8). For example, within the alpha channel (A) of one pixel, transparency represented by four bits and depth value represented by four bits can be concatenated. In the example, among the eight bits of the alpha channel (A), the most significant bit (MSB) and three bits adjacent to the MSB can represent transparency, and the remaining four bits can represent the depth value. For example, a sequence of bits (e.g., a bit sequence) representing a depth value can be located after a least significant bit (LSB) of a sequence of four bits representing transparency. Embodiments of the present disclosure are not limited thereto. For example, among the 8 bits of the alpha channel (A), the LSB and the 3 bits adjacent to the LSB can represent transparency, and the remaining 4 bits can represent the depth value.

[0060] For example, based on determining that depth values ​​are represented by a third number of bits (e.g., 7) that is greater than the first number of bits, the electronic device can use the depth information to generate an image that includes depth values ​​represented by the third number of bits, and transparencies represented by a fourth number of bits (e.g., 1 = 8 - 7) that is obtained by subtracting the third number of bits from the total number (e.g., 8). For example, if the image acquired based on operation (310) includes only completely transparent areas and completely opaque areas, the transparencies of all pixels can be represented by only two values, and thus the electronic device can represent the transparency using only one bit, and the depth value using the remaining seven bits. In the above example, within the alpha channel (A) of one pixel, a sequence of bits (e.g., a bit sequence) representing a depth value can be located after a bit representing transparencies. Embodiments of the present disclosure are not limited thereto. For example, within the alpha channel (A) of a pixel, a bit sequence representing a depth value may be positioned before the most significant bit (MSB) of a bit sequence (or bit(s)) representing transparency.

[0061] In one embodiment, since the bits representing the depth value are located in the portion containing the LSB within the alpha channel, the size of the concatenated value of the alpha channel may be related to the transparency among the transparency and depth values. For example, another electronic device that cannot obtain the depth value from the alpha channel may determine the concatenated value as transparency. When the concatenated value is determined as transparency, since the bits representing the depth value are located in the portion containing the LSB within the alpha channel, the order of the sizes of the transparencies of the pixels may match the order of the sizes of the transparencies included in the concatenated value, even though the concatenated value further includes the depth value.

[0062] In one embodiment, the electronic device can collectively determine the positions and / or sizes of transparency and depth values ​​within an alpha channel for all pixels, or independently for each pixel. When generating a file (110) representing a video, the electronic device can collectively set the positions and / or sizes of transparency and depth values ​​within an alpha channel for image frames included in the video.

[0063] Embodiments of the present disclosure are not limited thereto. For example, within image frames, the positions and / or sizes of transparency and depth values ​​within the alpha channel may be different from each other.

[0064] In one embodiment, an electronic device that generates pixel data representing pixels in which concatenated values ​​of transparency and depth values ​​are located in an alpha channel may generate or obtain metadata (e.g., metadata 120 of FIGS. 1A and / or 1B ) that includes information for extracting the transparency and the depth values ​​from the concatenated values. For example, the electronic device may generate metadata that indicates the number of bits within the alpha channel reserved for representing each of the depth values. In one embodiment, if the alpha channel has a designated number of bits (e.g., 8), the electronic device may generate metadata that indicates a ratio between the number of bits corresponding to the depth value and the number of bits corresponding to the transparency. For example, the electronic device may generate metadata that indicates the digits of one or more bits within the alpha channel occupied by the depth value. For example, the electronic device may generate metadata that indicates the number of bits within the alpha channel reserved for representing transparency and / or the digits of one or more bits representing the transparency. The electronic device can generate a file including the metadata and pixel data representing the image.

[0065] In one embodiment, the electronic device generates metadata indicating properties of concatenated values ​​of an alpha channel (e.g., location of transparency and / or depth information within the concatenated values, and / or size), although embodiments of the present disclosure are not limited thereto.

[0066] For example, since a completely transparent pixel (e.g., a pixel with maximum transparency) may not display any color, the electronic device may represent the property of the concatenated value using bits for representing color within the pixel (e.g., 24 bits representing a red channel, a green channel, and a blue channel). For example, when generating an image including a first region corresponding to a virtual object and a second region surrounding the first region, the transparencies of the pixels of the image corresponding to the second region may have maximum transparency (e.g., a binary number representing 100% transparency). In the example, the electronic device may add depth values ​​and transparencies to the alpha channel of the pixels corresponding to the first region (e.g., a concatenated value of the depth value and the transparency). In the example, the electronic device may add, to the pixels corresponding to the second region, bit numbers (or positions) of the depth values ​​and / or transparencies added to the alpha channel of the pixels corresponding to the first region.

[0067] In one embodiment, the range and step of the depth value expressed by the depth information can be adjusted. For example, if the maximum value of the depth value of a specific image is 10 and the depth value is expressed using 6 bits (e.g., bits representing natural numbers from 0 to 63), the electronic device can set the depth levels expressed by the bits in units of 10 / 64 = 0.156. By changing the unit of the depth level, the electronic device can generate a file (110) in which depth information determined in a range from 0 to 10 is expressed in detail.

[0068] Referring to FIG. 3A, when displaying an image and / or video of a virtual object represented by a file (110), the electronic device may perform at least one of operations (320, 322, 324). In operation (320), the electronic device may decode a depth value included in the file (110). For example, the electronic device may extract, identify, or parse a depth value from an alpha channel of pixels included in pixel data of the file (110). For example, the electronic device may segment concatenated values ​​included in the alpha channel of the pixels to obtain or identify depth values ​​corresponding to each of the pixels. In operation (322), the electronic device may perform depth rendering of the image and / or video using the decoded depth values. The depth rendering may include an operation of determining depth values ​​of each pixel of a two-dimensional image and / or a two-dimensional video of the file (110). Within the operation (324), the electronic device may perform three-dimensional rendering of the two-dimensional image and / or the two-dimensional video using depth values ​​determined based on depth rendering. The three-dimensional rendering may include an operation of displaying pixels having colors represented by a red channel, a green channel, and a blue channel according to a binocular parallax corresponding to a depth value obtained from a concatenated value of an alpha channel, based on a transparency obtained from the concatenated value.

[0069] In one embodiment, an electronic device can add depth values ​​to a file (110) without changing the data structure of the file (110) based on a red channel, a green channel, a blue channel, and an alpha channel. Since the data structure of the file (110) does not change, operations included in a pipeline for 3D rendering (e.g., an operation of identifying pixels based on the red channel, the green channel, and the blue channel) can be at least partially reused or maintained. Since the data structure does not change, the electronic device can generate or store a file (110) including more depth values ​​without increasing the size of the file (110).

[0070] FIG. 3B illustrates programs executed by an electronic device to create and / or execute a file (110) (e.g., display images and / or videos represented by the file (110). The programs executed by the electronic device may include an avatar data hub (351), a space flinger (352), a composition presentation manager (CPM) (354), an emoji studio (356), an avatar service (358), a camera service (359), an avatar camera hardware abstraction layer (HAL) (360), or any combination thereof. Data used to execute the programs (e.g., an avatar DB (355), and / or setting value(s) (357)) may be stored within the electronic device. By executing the above programs, the electronic device may generate or obtain a bitstream (e.g., IStream (361)) representing images and / or video, which may be displayed on a display or stored in a file (e.g., file (110) of FIG. 1A and / or FIG. 1B).

[0071] Referring to FIG. 3B, the avatar data hub (351) executed by the electronic device may be referred to as an avatar provider (or provider). The electronic device executing the avatar data hub (351) may perform two-dimensional rendering of an avatar (e.g., a virtual object (140) of FIG. 1A) in a two-dimensional buffer (e.g., a composite layer (353)). In one embodiment, to maintain compatibility, the avatar HAL (360) may be configured to perform rendering of the avatar. The electronic device may perform the rendering using the avatar HAL (360) to generate a bitstream (e.g., IStream (361)). The electronic device executing the avatar data hub (351) may obtain depth information about the avatar. The electronic device may obtain the depth information using information about the avatar stored in the avatar DB (355). The electronic device may obtain or generate pixel data (e.g., pixel data (130) of FIG. 1A and / or FIG. 1B) in which a concatenated value of a depth value and a transparency are included in an alpha channel by combining (e.g., combining based on quantization) the transparencies of pixels of a two-dimensional image of an avatar, represented by a two-dimensional buffer, and the depth values, represented by the depth information. The pixel data may be stored in an image buffer allocated to a memory (e.g., memory (220) of FIG. 2).

[0072] The avatar data hub (351) may include a stream interface for transmitting an image buffer to a virtual camera. Through the stream interface, the electronic device may generate (e.g., render) an image stream (IStream (361)) representing the avatar. The avatar data hub (351) may include a resource manager that manages information and / or resources used for rendering the avatar. The avatar data hub (351) may be configured to provide the resources, events related to the avatar, information tracked for rendering the avatar (e.g., information tracked by the sensor (240) of FIG. 2), and / or audio signals representing the sound of the avatar.

[0073] According to one embodiment, when an electronic device renders an avatar based on a file (e.g., file (110) of FIGS. 1A and / or 1B), the electronic device may generate or obtain an image and / or video of the avatar to be displayed on the display using a composite layer (353). For example, the space flinger (352) may generate the composite layer (353) using information about other layers managed by the CPM (354) and displayed by the electronic device. When the electronic device renders the avatar, the composite layer (353) may further include depth information and / or depth values. For example, the alpha channel of pixels in the composite layer (353) may include concatenated values ​​of transparency and depth values. Using the concatenated values, the electronic device may stereoscopically display a two-dimensional image and / or two-dimensional video of the avatar represented by the composite layer (353).

[0074] FIGS. 4A and 4B illustrate a flowchart of an electronic device according to one embodiment. The electronic device (101) of FIGS. 1A, 1B, and / or 2 may include the electronic device of FIGS. 4A and 4B. The operations of FIGS. 4A and 4B may be performed by the electronic device (101) and / or the processor (210) of FIG. 2. The order of the operations illustrated in FIGS. 4A and 4B is not limited to the order illustrated in FIGS. 4A and 4B. For example, the operations illustrated in FIGS. 4A and 4B may be performed in a different order than the order of the operations illustrated in FIGS. 4A and 4B. For example, at least two of the operations illustrated in FIGS. 4A and 4B may be performed substantially simultaneously.

[0075] Referring to FIG. 4A, in operation (410), according to one embodiment, an electronic device may perform rendering on a visual object (e.g., a virtual object (140) of FIG. 1A and / or FIG. 1B) to obtain a two-dimensional image and depth information for the visual object. For example, the electronic device may obtain or generate a two-dimensional image for a visual object placed in a three-dimensional virtual space. The two-dimensional image may represent the appearance (or external appearance) of the visual object as seen from a virtual camera placed in the virtual space. The two-dimensional image may represent the visual object and the appearance of the virtual space surrounding the visual object. The pixels of the two-dimensional image may represent the appearance of portions of the virtual space corresponding to each of the pixels, based on color and transparency. A two-dimensional image of an action (410) may be referred to as an RGBA (red-green-blue-alpha) image, in that it contains pixels based on red, green, blue, and transparency (e.g., transparency referred to as an alpha value).

[0076] The depth information of the operation (410) may represent depth values ​​of each of the pixels. For example, the depth values ​​may represent distances between parts of the virtual space corresponding to each pixel and a virtual camera. The depth information may be referred to as a depth map for the two-dimensional image of the operation (410).

[0077] Referring to FIG. 4A, in operation 420, an electronic device according to an embodiment may encode depth values ​​represented by depth information into an alpha channel of a two-dimensional image. The encoding may include an operation of converting (e.g., quantizing) transparency included in the alpha channel into a number of bits less than the number of bits of the alpha channel. The encoding may include an operation of obtaining a depth value expressed in a number of bits equal to the difference between the number of bits of the quantized transparency and the number of bits of the alpha channel. The encoding may include an operation of combining the quantized transparency and the obtained depth value to obtain a concatenation value of the transparency and the depth value. The encoding may include an operation of adding (or writing) the concatenation value to the alpha channel.

[0078] Referring to FIG. 4A, in operation (430), according to an embodiment, an electronic device may generate and / or store an image file (e.g., file (110) of FIG. 1A and / or FIG. 1B) representing a visual object, wherein depth values ​​and transparencies are included in an alpha channel. The electronic device may generate metadata including information for decoding (or parsing) depth values ​​and transparencies from the alpha channel of the pixels, together with pixel data including pixels having an alpha channel of operation (430) (e.g., pixel data (130) of FIG. 1A and / or FIG. 1B). The image file of operation (430) may include the pixel data and the metadata.

[0079] As described above with reference to FIG. 4A, according to one embodiment, an electronic device can generate an image file that further includes depth values ​​corresponding to each pixel without increasing the number of channels of the pixels (e.g., four channels including a red channel, a green channel, a blue channel, and an alpha channel). For example, the electronic device can change the use of the alpha channel to a channel in which all transparency and depth values ​​are embedded.

[0080] FIG. 4B illustrates an exemplary operation of an electronic device displaying an image and / or video contained in an image file of operation (430). The image file of operation (430) may be transmitted to another electronic device different from the electronic device that generated the image file. The other electronic device receiving the image file may perform the operation of FIG. 4B to decode and / or render (e.g., 2.5-dimensional rendering) the image file. Referring to FIG. 4B , within operation (450), according to an embodiment, the electronic device may identify an image file (e.g., the image file of operation (430)) representing a visual object (e.g., the visual object of operation (410)). The electronic device may perform operation (450) in response to an input for selecting or opening the image file. The electronic device may perform operation (450) in response to an input for browsing the image file. The above input may be detected or identified based on a tap gesture (or double-tap gesture) on an icon representing an image file, a mouse click (or mouse double-click), a gaze input, a hand gesture (e.g., a pinch gesture), and / or an utterance identified from an audio signal (e.g., "open that file").

[0081] Referring to FIG. 4B , in operation (460), according to one embodiment, the electronic device may decode values ​​encoded in an alpha channel of pixels of an image file to obtain depth values ​​and transparencies corresponding to each of the pixels. The value of the alpha channel may be a concatenated value in which transparency and depth values ​​are concatenated. For example, from the MSB of the concatenated value, the transparency and depth values ​​may be sequentially stored. The electronic device may obtain or identify information required for decoding in operation (460) from metadata of the image file (e.g., metadata (120) of FIG. 1A and / or FIG. 1B ). For example, from the metadata, the electronic device may obtain information for segmenting or parsing concatenated values ​​included in the alpha channel. The above information may include at least one of positions and / or locations within a concatenated value of bits representing transparency, positions and / or locations within a concatenated value of bits representing depth values.

[0082] Referring to FIG. 4B, in operation (470), according to an embodiment, the electronic device may generate a three-dimensional point corresponding to at least one of the pixels in the virtual space using the acquired depth values ​​and transparencies. For example, the electronic device may generate three-dimensional points each corresponding to pixels having a different transparency than a designated transparency representing a fully transparent pixel. The three-dimensional points may have three-dimensional coordinate values ​​in the virtual space based on the depth values ​​acquired based on operation (460). For example, the location of the three-dimensional point in the virtual space may be determined based on the location of the pixel corresponding to the three-dimensional point in the two-dimensional image and the depth value. When a plurality of three-dimensional points are generated, the electronic device may acquire or identify a point cloud representing a visual object including the plurality of three-dimensional points.

[0083] Referring to FIG. 4B, in operation (480), an electronic device according to one embodiment may perform rendering based on a virtual space including one or more three-dimensional points to obtain an image and / or video representing a visual object in three dimensions. For example, the electronic device may obtain or generate an image and / or video representing the appearance and / or appearance of three-dimensional points as viewed from a virtual camera within the virtual space, as defined by metadata of the image file.

[0084] Referring to FIG. 4B, in operation (490), according to one embodiment, the electronic device may display the acquired image and / or video. The electronic device may display the image and / or the video on the display. Operations (480, 490) may be referred to as three-dimensional rendering operations for a visual object. Based on the three-dimensional rendering, the electronic device may provide a three-dimensional representation of the visual object. For example, the electronic device may display an image and / or video having binocular parallax. For example, the electronic device may display an image and / or video representing a visual object that is rotated three-dimensionally according to a user's gesture.

[0085] FIG. 5 illustrates an exemplary operation of an electronic device performing scaling on a depth value. The electronic device (101) of FIG. 1A, FIG. 1B, and / or FIG. 2 may include the electronic device of FIG. 5. The electronic device (101) and / or the processor (210) of FIG. 2 may perform the operation of the electronic device described with reference to FIG. 5.

[0086] Fig. 5 illustrates exemplary virtual spaces (501, 502) obtained by performing rendering on a virtual object (510). When an electronic device generates a two-dimensional image representing a virtual object (510), the electronic device can determine depth values ​​corresponding to each pixel of the two-dimensional image based on a depth axis (d).

[0087] For example, an electronic device that creates a first virtual space (501) of FIG. 5 may determine a depth value corresponding to point p3 of a virtual object (510) as 10. The depth value may be mapped to a pixel corresponding to point p3 on a two-dimensional image representing the virtual object (510). Similarly, the electronic device may determine a depth value corresponding to point p4 of the virtual object (510) as 4. The depth value may be linked to a specific pixel of a two-dimensional image corresponding to point p4.

[0088] Referring to the exemplary first virtual space (501) of FIG. 5, although the depth value ranges from 0 to 128, since the virtual object (510) is placed on a portion of the first virtual space (501) having a depth value between 0 and 10, the depth values ​​corresponding to pixels representing the virtual object (510) may be determined only within the range between 0 and 10. For example, depth values ​​between 11 and 128 may not be used. According to one embodiment, the electronic device may perform scaling on the depth value based on the range of depth values ​​corresponding to pixels of the two-dimensional image.

[0089] For example, an electronic device that performs rendering for a virtual object (510) based on a first virtual space (501) can determine whether a range between a maximum value and a minimum value of depth values ​​corresponding to pixels representing the virtual object (510) is smaller than the entire range of depth values. If the range is smaller than the entire range (e.g., less than a specified percentage of the entire range), the electronic device can change the depth values ​​of the pixels by scaling the depth axis (d).

[0090] FIG. 5 illustrates a second virtual space (502) obtained by scaling the depth axis (d). Based on the second virtual space (502), the depth value of a pixel corresponding to a point p5 of a virtual object (520) (corresponding to a point p3 of a virtual object (510) in the first virtual space (501)) may be determined as 128. Based on the second virtual space (502), the depth value of a pixel corresponding to a point p6 of a virtual object (520) (corresponding to a point p4 of a virtual object (510) in the first virtual space (501)) may be determined as 10. For example, within the second virtual space (502), the depth values ​​of pixels representing the virtual object (520) may be determined in a range from 0 to 128. That is, by using the second virtual space (502) in which the depth axis (d) is scaled, the electronic device can obtain more precisely determined depth values. By performing encoding based on the above depth values, the electronic device can generate or store a file (e.g., file (110) of FIG. 1A and / or FIG. 1B) containing detailed depth values.

[0091] In one embodiment, an image obtained based on the operation of FIG. 5 may be used as part of a video representing the motion of a virtual object (520). Hereinafter, with reference to FIG. 6, an exemplary operation of an electronic device for generating a video including image frames representing the motion of a virtual object (520) and depth information corresponding to the image frames (e.g., depth information encoded in an alpha channel of pixels of the image frames) is described.

[0092] FIG. 6 illustrates exemplary operations of an electronic device for generating a video based on key frames (e.g., a first image (631) and / or a fifth image (635)). The electronic device (101) of FIG. 1A, FIG. 1B, and / or FIG. 2 may include the electronic device of FIG. 6. The electronic device (101) and / or the processor (210) of FIG. 2 may perform the operations of the electronic device described with reference to FIG. 6.

[0093] Figure 6 illustrates an exemplary operation of an electronic device that generates a video comprising a plurality of images representing the motion of a virtual object. The motion may represent a user's motion, as described below with reference to Figure 7. The video may be generated based on input indicating the creation and / or recording of a video.

[0094] Figure 6 illustrates exemplary images (631, 632, 633, 634, 635) acquired at consecutive time points (t1, t2, t3, t4, t5) in the time domain. The electronic device may generate pixel data (630) representing a sequence of the images (631, 632, 633, 634, 635). The pixel data (630) may include images (631, 632, 633, 634, 635) compressed based on a compression algorithm (e.g., a compression algorithm based on lossy compression).

[0095] In one embodiment, an electronic device that generates a video using images (631, 632, 633, 634, 635) may determine an image at a specific point in time as a reference image for other images after the specific point in time. The reference image may be referred to as a key frame. Fig. 6 illustrates an exemplary state in which a first image (631) corresponding to a point in time t1 is determined as a key frame. The electronic device may store, in an alpha channel of a first pixel having coordinates (x, y) of the first image (631), a concatenation value of transparency (AAAA) expressed in 4 bits and depth values ​​(DDDD) expressed in 4 bits.

[0096] In one embodiment, when an electronic device determines a specific image as a key frame, the electronic device may set, in the time domain, depth values ​​of pixels of one or more images included in a designated time interval after the specific image set as a key frame, as difference values ​​with respect to the depth values ​​of pixels of the specific image. Referring to FIG. 6, when a first image (631) corresponding to a time point t1 is determined as a key frame, the electronic device may set depth values ​​corresponding to pixels of images (632, 633, 634) of time points t2, t3, and t4 included in a designated time interval after the time point t1, as difference values ​​with respect to the depth values ​​of pixels of the first image (631).

[0097] For example, a difference value (D'D'D'D') between a depth value of a first pixel having coordinates (x, y) of a first image (631) and a depth value of the second pixel may be stored in an alpha channel of a second pixel having coordinates (x, y) of a second image (632) corresponding to a time point t2. In order to obtain the depth value of the second pixel, the electronic device may obtain depth information based on the shape of the virtual object at the time point t2. The electronic device may obtain difference values ​​between depth values ​​indicated by the depth information and depth values ​​of the first image (631). Similarly, a difference value (A'A'A'A') between transparency of the first pixel and transparency of the second pixel may be stored in an alpha channel of the second pixel.

[0098] Similarly, a difference value (D'D'D'D') between a depth value of a first pixel having coordinates (x, y) of a first image (631) and a depth value of the third pixel may be stored in an alpha channel of a third pixel having coordinates (x, y) of a third image (633) corresponding to time point t3. A difference value (A'A'A'A') between a transparency of the first pixel and a transparency of the third pixel may be stored in an alpha channel of the third pixel.

[0099] Similarly, in the alpha channel of the fourth pixel having the (x, y) coordinates of the fourth image (634) corresponding to the time point t4, the difference value for the transparency of the first pixel having the (x, y) coordinates of the first image (631) and the difference value for the depth value of the first pixel can be concatenated (A'A'A'A'D'D'D'D'). When performing three-dimensional rendering on the fourth image (634), the electronic device that identifies the concatenated value of the difference values ​​from the pixel data (630) can restore or obtain the depth values ​​and transparencies of the fourth image (634) by using the transparency and / or depth value of the first image (631), which is a key frame.

[0100] Referring to FIG. 6, the electronic device can determine the fifth image (635) of the time point t5 after the time point t4 as a key frame. In the alpha channel of the pixels of other images after the fifth image (635), the electronic device can store the difference values ​​for the depth values ​​and transparency of the pixels of the fifth image (635), respectively. (A2A2A2A2D2D2D2D2)

[0101] Referring to FIG. 6, an electronic device may generate pixel data (630) representing a video including images (631, 632, 633, 634, 635). The electronic device may generate or store a file (610) including the pixel data (630) and metadata (620). The electronic device may store, within the metadata (620), information representing one or more key frames (e.g., the first image (631) and / or the fifth image (635)), and information regarding a concatenation value stored in an alpha channel of other frames between the key frames.

[0102] As described above, in one embodiment, the electronic device selects key frames (e.g., the first image (631) and / or the fifth image (635)) based on a specified time interval and calculates values ​​to be stored in the alpha channels of other frames, but embodiments of the present disclosure are not limited thereto.

[0103] For example, after determining the first image (631) of time point t1 as a key frame, the electronic device may store, in the alpha channels of a specified number of images positioned after the first image (631) in the time domain, difference values ​​for depth values ​​of the first image (631). For example, when the specified number is 3, the electronic device may store, in the alpha channels of three images (e.g., a second image (632), a third image (633), and a fourth image (634)) after the first image (631), difference values ​​for depth values ​​of pixels of each of the images with respect to depth values ​​of pixels of the first image (631).

[0104] In one embodiment, when generating a video representing continuous motion of a virtual object, images included in the video may have relatively small differences unless abrupt motion occurs. For example, the difference value between the depth values ​​of a first image (631) set as a key frame and other images may be determined within a relatively small numerical range. Based on the numerical range, the electronic device may reduce the size (e.g., number of bits) of the alpha channel of the other images and / or the depth values ​​encoded in the alpha channel. The reduction in size may reduce the size of pixel data (630) and / or file (610) representing the video.

[0105] In one embodiment, when generating a video representing continuous motion of a virtual object, when a rapid motion occurs, images included in the video may have a relatively large difference. In this case, the differences in depth values ​​of the images may change within a relatively large numerical range. According to one embodiment, when the difference between the depth values ​​of the first image (631), which is a key frame, and another image exceeds a reference range, the electronic device may store the depth value of the other image instead of storing the difference value between the depth value of the first image (631) and the depth value of the other image in the alpha channel of the other image. For example, when the difference between the depth values ​​of the first image (631), which is a key frame, and the other image exceeds a reference range, the concatenated values ​​of the depth values ​​and transparencies of the pixels of the other image may be included in the alpha channel of the pixels of the other image.

[0106] FIG. 7 illustrates exemplary operations of an electronic device (101) that generates a video representing the motion of a visual object (720) using sensor data. The electronic device (101) of FIG. 1A, FIG. 1B, and / or FIG. 2 may include the electronic device (101) of FIG. 7. The electronic device (101) and / or the processor (210) of FIG. 2 may perform the operations of the electronic device (101) described with reference to FIG. 7. The operations described with reference to FIG. 7 may be related to at least one of the operations of FIG. 4A.

[0107] FIG. 7 illustrates an electronic device (101) in the form of an HMD, including a plurality of displays (711, 712). Each of the plurality of displays (711, 712) may be configured to be positioned toward both eyes of the user (105) when worn by the user (105). For example, the first display (711) may be positioned toward the left eye of the user (105), and the second display (712) may be positioned toward the right eye of the user (105). A combination of the plurality of displays (711, 712) may be referred to as a display assembly and / or a display module.

[0108] Referring to FIG. 7, when the visual object (720) is set to simulate the motion of the user (105), the electronic device (101) can obtain sensor data representing the motion of the user (105) by using a sensor configured to detect the motion of the user (105) (e.g., sensor (240) of FIG. 2). For example, the electronic device (101) can obtain sensor data representing the direction (d_hmd) of the electronic device (101) by using the sensor. When the user (105) wears the electronic device (101), the direction (d_hmd) of the electronic device (101) indicated by the sensor data can be linked to the direction of the head of the user (105). When the visual object (720) is set to mimic the motion of the user (105), the direction (d_avt) of the visual object (720) shown through the displays (711, 712) can be changed synchronously with the direction (d_hmd) of the electronic device (101).

[0109] FIG. 7 illustrates an exemplary state of an electronic device (101) generating a video representing the motion of a visual object (720), referred to as an avatar and / or virtual object. The electronic device (101) may initiate generation of the video based on a specified input. If the visual object (720) is configured to mimic the motion of a user (105), the video generated by the electronic device (101) may represent the visual object (720) mimicking the motion of the user (105) detected during video generation.

[0110] Referring to FIG. 7, while generating (or recording) a video, the electronic device (101) may display a visual object (730) on the displays (711, 712) for receiving an input to stop generating the video. Although FIG. 7 illustrates a visual object (730) including a designated text such as “stop,” the shape or location of the visual object (730) of the present disclosure is not limited to the visual object (730) of FIG. 7. While generating the video, the electronic device (101) may detect a motion of the user (105) using sensor data detected from a sensor. While generating the video, the electronic device (101) may at least partially change the visual object (720) displayed on the displays (711, 712) to mimic the detected motion. While generating the video, the electronic device (101) may acquire a plurality of images representing the at least partially changed visual object (720). The above multiple images may be included in one file (e.g., file (610) of FIG. 6) based on the operation described with reference to FIG. 6.

[0111] According to one embodiment, the electronic device (101) may select a specific image as a key frame from among images representing the motion of the visual object (720), as described above with reference to FIG. 6. The depth values, transparency, and / or colors of pixels of other images may be expressed as difference values ​​for the depth values, transparency, and / or colors of pixels of the specific image selected as the key frame. The electronic device (101) may select a key frame based on the intensity (and / or size) of the motion of the user (105) indicated by the sensor data.

[0112] For example, if a first image of a first time point is selected as a key frame, the electronic device (101) can detect sensor data representing the motion of the user (105) from a sensor at a second time point after the first time point. The electronic device (101) can identify a difference between the sensor data detected at the first time point and the sensor data detected at the second time point. If the difference is included in a reference range, the electronic device can store difference values ​​of depth values ​​of the first image and the second image in the alpha channel of the pixels of the second image. For example, the electronic device can generate a second image corresponding to the second time point, in which difference values ​​of depth values ​​included in the alpha channel of the first image and other depth values ​​represented by other depth information acquired based on the sensor data of the second time point are each included in the alpha channel of the pixels.

[0113] For example, if the difference between the sensor data of the first time point and the sensor data of the second time point is outside the reference range (e.g., if the difference exceeds the reference range), the electronic device may store the concatenated values ​​of the depth values ​​of the second image and the transparencies of the second image in the alpha channel of the pixels of the second image. For example, the electronic device may generate a second image corresponding to the second time point, in which the different depth values ​​are each included in the alpha channel of the pixels. The electronic device may generate a file including the first image and the second image.

[0114] As described above, while generating a video for a visual object (720) that mimics the motion of a user (105), the electronic device (101) may select or determine a key frame based on the intensity and / or size of the motion. For example, when detecting a relatively rapid motion, the electronic device (101) may determine an image acquired at that time as a key frame. When detecting a relatively small motion, the electronic device (101) may select or determine a key frame based on criteria (e.g., a specified period and / or a specified number) described with reference to FIG. 6.

[0115] FIG. 8 illustrates an exemplary operation of an electronic device (101) that performs three-dimensional rendering on a visual object (810) represented by an image file. The electronic device (101) of FIG. 1A, FIG. 1B, and / or FIG. 2 may include the electronic device (101) of FIG. 8. The electronic device (101) and / or the processor (210) of FIG. 2 may perform the operation of the electronic device (101) described with reference to FIG. 8. The operation described with reference to FIG. 8 may be related to at least one of the operations of FIG. 4B.

[0116] FIG. 8 illustrates an exemplary state of an electronic device (101) performing rendering for a visual object (810) referred to as an avatar and / or virtual object. The electronic device (101) may receive an input for displaying an image and / or a video represented by a file (e.g., file (110) of FIG. 1A and / or FIG. 1B and / or file (610) of FIG. 6). The electronic device (101) receiving the input may obtain colors, transparencies, and depth values ​​of pixels of the image and / or the video from pixel data of the file. For example, the electronic device (101) may identify or obtain depth values ​​included in an alpha channel of the pixels. Based on the obtained depth values, the electronic device (101) may determine binocular disparity of each of the pixels. The electronic device (101) can display an image representing a visual object (810) on the first display (711) based on the determined binocular disparity. The electronic device (101) can display another image representing a visual object (810) shifted based on the binocular disparity on the second display (712).

[0117] For example, according to the depth values ​​identified from the concatenated values ​​of the alpha channel, portions of the visual object (810) displayed on the first display (711) may be shifted from portions of the visual object (810) displayed on the second display (712), respectively. Referring to FIG. 8, when displaying a visual object (810) in the form of an avatar wearing a hat, the binocular disparity of a portion of the visual object (810) that is set to be relatively close to the user (105) (e.g., the portion corresponding to the hat (811)) may be greater than that of another portion of the visual object (810) (e.g., the portion corresponding to the hair (812)). From the alpha channel of the pixels of the image included in the file, the electronic device (101) may identify the depth values ​​and transparencies of portions of the visual object (810) corresponding to each of the pixels. The electronic device (101) can control a display assembly including displays (711, 712) to display a visual object (810) expressed based on the above transparencies. Referring to FIG. 8, by the pixels having the transparencies, a background area (820) beyond the visual object (810) can be displayed in a portion of the display area adjacent to the visual object (810) (e.g., a portion outside the boundary of the face expressed by the visual object (810).

[0118] In one embodiment, when playing a video representing the motion of a visual object (810), the electronic device (101) may display a visual object (830) on the displays (711, 712) to at least temporarily stop the playback of the video. The video may be included in a file (e.g., file (610) of FIG. 6) generated based on the operations described with reference to FIG. 6 and / or FIG. 7.

[0119] As described above, according to one embodiment, the electronic device (101) may provide information for three-dimensional rendering (e.g., point cloud rendering) of a visual object (810), such as an avatar, together with a two-dimensional image of the visual object (810), by adding depth values ​​as well as transparency to an alpha channel. A file generated by the electronic device (101) (e.g., file (110) of FIGS. 1A and / or 1B and / or file (610) of FIG. 6 ) may include pixels based on a red channel, a green channel, a blue channel, and an alpha channel, and since no additional channel for storing depth values ​​is defined, the file may be compatible with a graphics pipeline that can read only the four channels described above (e.g., hardware, software, or a combination thereof for expressing color and transparency excluding depth values ​​using the four channels). For example, an external electronic device executing a conventional graphics pipeline may perform two-dimensional rendering on a visual object (810) to generate or display a two-dimensional image and / or two-dimensional video of the visual object (810).

[0120] FIG. 9A illustrates an example of a perspective view of an electronic device, according to one embodiment. FIG. 9B illustrates an example of one or more hardware components disposed within an electronic device (101). According to one embodiment, the electronic device (101) may have a form of glasses that are wearable on a body part (e.g., head) of a user (e.g., user 105 of FIG. 1A and / or FIG. 1B). The electronic device (101) of FIGS. 9A and 9B may be an example of the electronic device (101) of FIG. 1A and / or FIG. 1B. The electronic device (101) may include a head-mounted display (HMD). For example, the housing of the electronic device (101) may include a flexible material, such as rubber and / or silicone, that is configured to fit closely to a portion of the user's head (e.g., a portion of the face surrounding both eyes). For example, the housing of the electronic device (101) may include one or more straps capable of being twined around the user's head, and / or one or more temples attachable to the ears of the head.

[0121] Referring to FIG. 9A, according to one embodiment, an electronic device (101) may include at least one display (950) and a frame (900) supporting at least one display (950).

[0122] According to one embodiment, the electronic device (101) can be worn on a part of a user's body. The electronic device (101) can provide augmented reality (AR), virtual reality (VR), or mixed reality (MR) that combines augmented reality and virtual reality to the user wearing the electronic device (101). For example, the electronic device (101) can display a virtual reality image provided from at least one optical device (982, 984) of FIG. 9B on at least one display (950) in response to a user's designated gesture acquired through the motion recognition cameras (960-2, 960-3) of FIG. 9B.

[0123] According to one embodiment, at least one display (950) may provide visual information to a user. For example, at least one display (950) may include a transparent or translucent lens. At least one display (950) may include a first display (950-1) and / or a second display (950-2) spaced apart from the first display (950-1). For example, the first display (950-1) and the second display (950-2) may be positioned at positions corresponding to the user's left and right eyes, respectively.

[0124] Referring to FIG. 9B, at least one display (950) can provide a user with visual information transmitted from external light and other visual information distinct from the visual information through a lens included in the at least one display (950). The lens can be formed based on at least one of a Fresnel lens, a pancake lens, or a multi-channel lens. For example, at least one display (950) can include a first surface (931) and a second surface (932) opposite to the first surface (931). A display area can be formed on the second surface (932) of the at least one display (950). When a user wears the electronic device (101), external light can be transmitted to the user by being incident on the first surface (931) and transmitted through the second surface (932). As another example, at least one display (950) can display an augmented reality image combined with a virtual reality image provided from at least one optical device (982, 984) on a real screen transmitted through external light, in a display area formed on the second surface (932).

[0125] In one embodiment, at least one display (950) may include at least one waveguide (933, 934) that diffracts light emitted from at least one optical device (982, 984) and transmits the diffracted light to a user. The at least one waveguide (933, 934) may be formed based on at least one of glass, plastic, or polymer. A nano-pattern may be formed on at least a portion of the exterior or interior of the at least one waveguide (933, 934). The nano-pattern may be formed based on a grating structure having a polygonal and / or curved shape. Light incident on one end of the at least one waveguide (933, 934) may be propagated to the other end of the at least one waveguide (933, 934) by the nano-pattern. At least one waveguide (933, 934) may include at least one diffractive element (e.g., a diffractive optical element (DOE), a holographic optical element (HOE)) and at least one reflective element (e.g., a reflective mirror). For example, at least one waveguide (933, 934) may be arranged within the electronic device (101) to guide a screen displayed by at least one display (950) to the user's eyes. For example, the screen may be transmitted to the user's eyes based on total internal reflection (TIR) ​​occurring within the at least one waveguide (933, 934).

[0126] The electronic device (101) can analyze an object included in a real image collected through a shooting camera (960-4), combine a virtual object corresponding to an object to be provided with augmented reality among the analyzed objects, and display the virtual object on at least one display (950). The virtual object can include at least one of text and an image regarding various information related to the object included in the real image. The electronic device (101) can analyze the object based on a multi-camera such as a stereo camera. For the object analysis, the electronic device (101) can perform spatial recognition (e.g., simultaneous localization and mapping (SLAM)) using the multi-camera and / or time-of-flight (ToF). A user wearing the electronic device (101) can view an image displayed on at least one display (950).

[0127] According to one embodiment, the frame (900) may be formed as a physical structure that allows the electronic device (101) to be worn on the user's body. According to one embodiment, the frame (900) may be configured so that, when the user wears the electronic device (101), the first display (950-1) and the second display (950-2) can be positioned corresponding to the user's left and right eyes. The frame (900) may support at least one display (950). For example, the frame (900) may support the first display (950-1) and the second display (950-2) to be positioned corresponding to the user's left and right eyes.

[0128] Referring to FIG. 9A, the frame (900) may include a region (920) that at least partially contacts a part of the user's body when the user wears the electronic device (101). For example, the region (920) of the frame (900) that contacts a part of the user's body may include a region that contacts a part of the user's nose, a part of the user's ear, and a part of the side of the user's face that the electronic device (101) makes contact with. According to one embodiment, the frame (900) may include a nose pad (910) that contacts a part of the user's body. When the electronic device (101) is worn by the user, the nose pad (910) may contact a part of the user's nose. The frame (900) may include a first temple (904) and a second temple (905) that contact a part of the user's body that is distinct from the part of the user's body.

[0129] For example, the frame (900) may include a first rim (901) that surrounds at least a portion of the first display (950-1), a second rim (902) that surrounds at least a portion of the second display (950-2), a bridge (903) that is disposed between the first rim (901) and the second rim (902), a first pad (911) that is disposed along a portion of the edge of the first rim (901) from one end of the bridge (903), a second pad (912) that is disposed along a portion of the edge of the second rim (902) from the other end of the bridge (903), a first temple (904) that extends from the first rim (901) and is fixed to a portion of an ear of the wearer, and a second temple (905) that extends from the second rim (902) and is fixed to a portion of an ear opposite the ear. The first pad (911) and the second pad (912) may be in contact with a portion of the user's nose, and the first temple (904) and the second temple (905) may be in contact with a portion of the user's face and a portion of the user's ear. The temples (904, 905) may be rotatably connected to the rim through the hinge units (906, 907) of FIG. 9B. The first temple (904) may be rotatably connected to the first rim (901) through the first hinge unit (906) disposed between the first rim (901) and the first temple (904). The second temple (905) may be rotatably connected to the second rim (902) through the second hinge unit (907) disposed between the second rim (902) and the second temple (905). According to one embodiment, the electronic device (101) can identify an external object (e.g., a user's fingertip) touching the frame (900) and / or a gesture performed by the external object by using a touch sensor, a grip sensor, and / or a proximity sensor formed on at least a portion of a surface of the frame (900).

[0130] According to one embodiment, the electronic device (101) may include hardwares that perform various functions (e.g., hardwares to be described later based on the block diagram of FIG. 11). For example, the hardwares may include a battery module (970), an antenna module (975), at least one optical device (982, 984), speakers (e.g., speakers 955-1, 955-2), microphones (e.g., microphones 965-1, 965-2, 965-3), a light-emitting module, and / or a printed circuit board (PCB) (990) (e.g., a printed circuit board). The various hardwares may be arranged within the frame (900).

[0131] According to one embodiment, a microphone (e.g., microphones 965-1, 965-2, 965-3) of the electronic device (101) may be disposed on at least a portion of the frame (900) to acquire a sound signal. A first microphone (965-1) disposed on the bridge (903), a second microphone (965-2) disposed on the second rim (902), and a third microphone (965-3) disposed on the first rim (901) are illustrated in FIG. 9B , but the number and arrangement of the microphones (965) are not limited to the embodiment of FIG. 9B . When the number of microphones (965) included in the electronic device (101) is two or more, the electronic device (101) may identify a direction of a sound signal by using a plurality of microphones disposed on different portions of the frame (900).

[0132] According to one embodiment, at least one optical device (982, 984) can project a virtual object onto at least one display (950) to provide various image information to a user. For example, at least one optical device (982, 984) can be a projector. At least one optical device (982, 984) can be disposed adjacent to at least one display (950) or can be included within at least one display (950) as a part of at least one display (950). According to one embodiment, the electronic device (101) can include a first optical device (982) corresponding to a first display (950-1) and a second optical device (984) corresponding to a second display (950-2). For example, at least one optical device (982, 984) may include a first optical device (982) disposed at an edge of a first display (950-1) and a second optical device (984) disposed at an edge of a second display (950-2). The first optical device (982) may transmit light to a first waveguide (933) disposed on the first display (950-1), and the second optical device (984) may transmit light to a second waveguide (934) disposed on the second display (950-2).

[0133] In one embodiment, the camera (960) may include a recording camera (960-4), an eye tracking camera (ET CAM) (960-1), and / or a motion recognition camera (960-2, 960-3). The recording camera (960-4), the eye tracking camera (960-1), and the motion recognition cameras (960-2, 960-3) may be positioned at different locations on the frame (900) and may perform different functions. The eye tracking camera (960-1) may output data indicating the position or gaze of the eyes of a user wearing the electronic device (101). For example, the electronic device (101) may detect the gaze from an image including the user's pupils obtained through the eye tracking camera (960-1). The electronic device (101) can identify an object (e.g., a real object and / or a virtual object) focused on by the user using the user's gaze acquired through the gaze tracking camera (960-1). The electronic device (101) that has identified the focused object can execute a function (e.g., gaze interaction) for interaction between the user and the focused object. The electronic device (101) can express a part corresponding to the eye of an avatar representing the user in a virtual space using the user's gaze acquired through the gaze tracking camera (960-1). The electronic device (101) can render an image (or screen) displayed on at least one display (950) based on the position of the user's eyes. For example, the visual quality of a first area related to the gaze within the image and the visual quality (e.g., resolution, brightness, saturation, grayscale, PPI) of a second area distinguished from the first area may be different from each other. The electronic device (101) can obtain an image having visual quality of a first area matching the user's gaze and visual quality of a second area using foveated rendering.For example, if the electronic device (101) supports an iris recognition function, user authentication can be performed based on iris information acquired using the gaze tracking camera (960-1). An example in which the gaze tracking camera (960-1) is positioned toward the user's right eye is illustrated in FIG. 9B, but the embodiment is not limited thereto, and the gaze tracking camera (960-1) may be positioned solely toward the user's left eye, or toward both eyes.

[0134] In one embodiment, the capturing camera (960-4) can capture an actual image or background to be aligned with a virtual image to implement augmented reality or mixed reality content. The capturing camera (960-4) can be used to acquire a high-resolution image based on HR (high resolution) or PV (photo video). The capturing camera (960-4) can capture an image of a specific object existing at a location viewed by the user and provide the image to at least one display (950). The at least one display (950) can display a single image in which information about an actual image or background including the image of the specific object acquired using the capturing camera (960-4) and a virtual image provided through at least one optical device (982, 984) are superimposed. The electronic device (101) can compensate for depth information (e.g., the distance between the electronic device (101) and an external object acquired through a depth sensor) using the image acquired through the capturing camera (960-4). The electronic device (101) can perform object recognition through an image acquired using the photographing camera (960-4). The electronic device (101) can perform a function of focusing on an object (or subject) in an image (e.g., auto focus) and / or an optical image stabilization (OIS) function (e.g., anti-shake function) using the photographing camera (960-4). The electronic device (101) can perform a pass-through function to display an image acquired through the photographing camera (960-4) by overlapping at least a portion of a screen representing a virtual space on at least one display (950). In one embodiment, the photographing camera (960-4) can be disposed on a bridge (903) disposed between the first rim (901) and the second rim (902).

[0135] The gaze tracking camera (960-1) can implement more realistic augmented reality by tracking the gaze of a user wearing the electronic device (101) and matching the user's gaze with visual information provided to at least one display (950). For example, when the electronic device (101) looks straight ahead, the electronic device (101) can naturally display environmental information related to the user's front at a location where the user is located on at least one display (950). The gaze tracking camera (960-1) can be configured to capture an image of the user's pupil to determine the user's gaze. For example, the gaze tracking camera (960-1) can receive gaze detection light reflected from the user's pupil and track the user's gaze based on the position and movement of the received gaze detection light. In one embodiment, the gaze tracking camera (960-1) can be positioned at positions corresponding to the user's left and right eyes. For example, the gaze tracking camera (960-1) may be positioned within the first rim (901) and / or the second rim (902) to face the direction in which the user wearing the electronic device (101) is positioned.

[0136] The gesture recognition camera (960-2, 960-3) can recognize the movement of the user's entire body, such as the user's torso, hand, or face, or a part of the body, and thereby provide a specific event on a screen provided on at least one display (950). The gesture recognition camera (960-2, 960-3) can recognize the user's gesture (gesture recognition), obtain a signal corresponding to the gesture, and provide a display corresponding to the signal on at least one display (950). The processor can identify the signal corresponding to the gesture, and perform a designated function based on the identification. The gesture recognition camera (960-2, 960-3) can be used to perform a spatial recognition function using SLAM and / or a depth map for 6 degrees of freedom pose (6 dof pose). The processor can perform a gesture recognition function and / or an object tracking function using the gesture recognition camera (960-2, 960-3). In one embodiment, the motion recognition cameras (960-2, 960-3) may be positioned on the first rim (901) and / or the second rim (902).

[0137] The camera (960) included in the electronic device (101) is not limited to the above-described gaze tracking camera (960-1) and motion recognition cameras (960-2, 960-3). For example, the electronic device (101) may identify an external object included in the FoV using a camera positioned toward the user's FoV. The electronic device (101) may identify an external object based on a sensor for identifying the distance between the electronic device (101) and the external object, such as a depth sensor and / or a time of flight (ToF) sensor. The camera (960) positioned toward the FoV may support an autofocus function and / or an optical image stabilization (OIS) function. For example, the electronic device (101) may include a camera (960) (e.g., a face tracking (FT) camera) positioned toward the face to obtain an image including the face of a user wearing the electronic device (101).

[0138] In one embodiment, the electronic device (101) may further include a light source (e.g., an LED) that emits light toward a subject (e.g., a user's eyes, face, and / or an external object within the FoV) being photographed using the camera (960). The light source may include an infrared wavelength LED. The light source may be disposed on at least one of the frame (900) and the hinge units (906, 907).

[0139] According to one embodiment, the battery module (970) may supply power to electronic components of the electronic device (101). In one embodiment, the battery module (970) may be disposed within the first temple (904) and / or the second temple (905). For example, the battery module (970) may be a plurality of battery modules (970). The plurality of battery modules (970) may be disposed within each of the first temple (904) and the second temple (905). In one embodiment, the battery module (970) may be disposed at an end of the first temple (904) and / or the second temple (905).

[0140] The antenna module (975) can transmit signals or power to the outside of the electronic device (101), or receive signals or power from the outside. In one embodiment, the antenna module (975) can be positioned within the first temple (904) and / or the second temple (905). For example, the antenna module (975) can be positioned close to one surface of the first temple (904) and / or the second temple (905).

[0141] The speaker (955) can output an acoustic signal to the outside of the electronic device (101). The acoustic output module may be referred to as a speaker. In one embodiment, the speaker (955) may be positioned within the first temple (904) and / or the second temple (905) so as to be positioned adjacent to the ear of a user wearing the electronic device (101). For example, the speaker (955) may include a second speaker (955-2) positioned within the first temple (904) and thus positioned adjacent to the user's left ear, and a first speaker (955-1) positioned within the second temple (905) and thus positioned adjacent to the user's right ear.

[0142] The light-emitting module may include at least one light-emitting element. The light-emitting module may emit light of a color corresponding to a specific state or emit light with an action corresponding to a specific state to visually provide a user with information regarding a specific state of the electronic device (101). For example, when the electronic device (101) requires charging, the electronic device (101) may emit red light at a regular cycle. In one embodiment, the light-emitting module may be disposed on the first rim (901) and / or the second rim (902).

[0143] Referring to FIG. 9B, according to one embodiment, an electronic device (101) may include a printed circuit board (PCB) (990). The PCB (990) may be included in at least one of the first temple (904) or the second temple (905). The PCB (990) may include an interposer disposed between at least two sub-PCBs. One or more hardwares included in the electronic device (101) (e.g., hardwares illustrated by different blocks in FIG. 11) may be disposed on the PCB (990). The electronic device (101) may include a flexible PCB (FPCB) for interconnecting the hardwares.

[0144] According to one embodiment, the electronic device (101) may include at least one of a gyro sensor, a gravity sensor, and / or an acceleration sensor for detecting a posture of the electronic device (101) and / or a posture of a body part (e.g., a head) of a user wearing the electronic device (101). Each of the gravity sensor and the acceleration sensor may measure gravitational acceleration and / or acceleration based on mutually perpendicular designated three-dimensional axes (e.g., an x-axis, a y-axis, and a z-axis). The gyro sensor may measure an angular velocity of each of the designated three-dimensional axes (e.g., an x-axis, a y-axis, and a z-axis). At least one of the gravity sensor, the acceleration sensor, and the gyro sensor may be referred to as an inertial measurement unit (IMU). According to one embodiment, the electronic device (101) may identify a motion and / or gesture of the user performed to execute or stop a specific function of the electronic device (101) based on the IMU.

[0145] FIGS. 10A and 10B illustrate an example of an exterior appearance of an electronic device (e.g., electronic device (101)). The electronic device (101) of FIGS. 10A and 10B may be an example of the electronic device (101) of FIGS. 1A and / or 1B. According to one embodiment, an example of an exterior appearance of a first side (1010) of a housing of the electronic device (101) is illustrated in FIG. 10A, and an example of an exterior appearance of a second side (1020) opposite to the first side (1010) may be illustrated in FIG. 10B.

[0146] Referring to FIG. 10A, according to one embodiment, a first surface (1010) of an electronic device (101) may have a form attachable on a body part of a user (e.g., the face of the user). In one embodiment, the electronic device (101) may further include a strap for fixing on a body part of a user, and / or one or more temples (e.g., the first temple (904) and / or the second temple (905) of FIGS. 9A and 9B). A first display (950-1) for outputting an image to a left eye among the user's two eyes, and a second display (950-2) for outputting an image to a right eye among the user's two eyes may be disposed on the first surface (1010). The electronic device (101) is formed on the first surface (1010) and may further include a rubber or silicone packing to prevent interference by light (e.g., ambient light) different from the light radiated from the first display (950-1) and the second display (950-2).

[0147] According to one embodiment, the electronic device (101) may include cameras (960-1) for photographing and / or tracking both eyes of the user adjacent to each of the first display (950-1) and the second display (950-2). The cameras (960-1) may be referred to as the gaze tracking camera (960-1) of FIG. 9B. According to one embodiment, the electronic device (101) may include cameras (960-5, 960-6) for photographing and / or recognizing the face of the user. The cameras (960-5, 960-6) may be referred to as FT cameras. The electronic device (101) may control an avatar representing the user in a virtual space based on the motion of the user's face identified using the cameras (960-5, 960-6). For example, the electronic device (101) may change the texture and / or shape of a portion of an avatar (e.g., a portion of an avatar representing a human face) using information obtained by cameras (960-5, 960-6) (e.g., FT cameras) and representing a facial expression of a user wearing the electronic device (101).

[0148] Referring to FIG. 10B, a camera (e.g., cameras (960-7, 960-8, 960-9, 960-10, 960-11, 960-12)) and / or a sensor (e.g., a depth sensor (1030)) for obtaining information related to the external environment of the electronic device (101) may be disposed on a second surface (1020) opposite to the first surface (1010) of FIG. 10A. For example, the cameras (960-7, 960-8, 960-9, 960-10) may be disposed on the second surface (1020) to recognize external objects. Cameras (960-7, 960-8, 960-9, 960-10) may be referenced to the motion recognition cameras (960-2, 960-3) of FIG. 9b.

[0149] For example, using cameras (960-11, 960-12), the electronic device (101) can obtain images and / or videos to be transmitted to each of the user's eyes. The camera (960-11) can be placed on the second surface (1020) of the electronic device (101) to obtain an image to be displayed through the second display (950-2) corresponding to the right eye among the two eyes. The camera (960-12) can be placed on the second surface (1020) of the electronic device (101) to obtain an image to be displayed through the first display (950-1) corresponding to the left eye among the two eyes. The cameras (960-11, 960-12) can be referred to as the shooting camera (960-4) of FIG. 9B.

[0150] According to one embodiment, the electronic device (101) may include a depth sensor (1030) disposed on the second face (1020) to identify a distance between the electronic device (101) and an external object. Using the depth sensor (1030), the electronic device (101) may obtain spatial information (e.g., a depth map) for at least a portion of the FoV of a user wearing the electronic device (101). In one embodiment, a microphone may be disposed on the second face (1020) of the electronic device (101) to obtain a sound output from an external object. The number of microphones may be one or more, depending on the embodiment.

[0151] Hereinafter, with reference to FIG. 11, the hardware or software configuration of the electronic device (101) is described.

[0152] Fig. 11 illustrates an example of a block diagram of an electronic device (e.g., an electronic device (101)). The electronic device (101) of Fig. 11 may be an example of the electronic device (101) of Fig. 1a and / or Fig. 1b, or the electronic device (101) of Figs. 9a to 10b.

[0153] Referring to FIG. 11, an electronic device (101) according to one embodiment may include a processor (1110), a memory (1115), a display (230) (e.g., the display (230) of FIG. 2, the first display (950-1) and / or the second display (950-2) of FIGS. 9A, 9B, 10A, and 10B), and / or a sensor (1120). The processor (1110), the memory (1115), the display (230), and / or the sensor (1120) may be electrically and / or operatively connected to each other by electronic components such as a communication bus (1102). The processor (1110), the display (230), and the memory (1115) of FIG. 11 may correspond to the processor (210), the display (230), and the memory (220) of FIG. 2, respectively. Among the descriptions of the processor (1110), display (230), and memory (1115) of FIG. 11, descriptions that overlap with the descriptions of the processor (210), display (230), and memory (220) of FIG. 2 may be omitted. The camera (1125) of FIG. 11 may correspond to the sensor (240) and / or image sensor of FIG. 2.

[0154] According to one embodiment, one or more instructions (or commands) representing data to be processed, calculations to be performed, and / or operations to be performed by the processor (1110) of the electronic device (101) may be stored in the memory (1115) of the electronic device (101). A set of one or more instructions may be referred to as a program, firmware, an operating system, a process, a routine, a sub-routine, and / or a software application (hereinafter, “application”). For example, the electronic device (101) and / or the processor (1110) may perform at least one of the operations of FIG. 4A and / or FIG. 4B when a set of a plurality of instructions distributed in the form of an operating system, firmware, a driver, a program, and / or a software application is executed. Hereinafter, the fact that a software application is installed in an electronic device (101) may mean that one or more instructions provided in the form of a software application (or package) are stored in a memory (1115), and that the one or more applications are stored in a format executable by a processor (1110) (e.g., a file having an extension specified by an operating system of the electronic device (101)). As an example, an application may include a program and / or a library related to a service provided to a user.

[0155] Referring to FIG. 11, programs installed in the electronic device (101) may be included in any one of different layers, including an application layer (1140), a framework layer (1150), and / or a hardware abstraction layer (HAL) (1180), based on the target. For example, programs (e.g., modules or drivers) designed to target the hardware (e.g., the display (230), and / or the sensor (1120)) of the electronic device (101) may be included in the hardware abstraction layer (1180). The framework layer (1150) may be referred to as an XR framework layer from the perspective of including one or more programs for providing an XR (extended reality) service. For example, the layers illustrated in FIG. 11 may be logically separated, and this does not mean that the address space of the memory (1115) is separated by the layers.

[0156] For example, within the framework layer (1150), programs designed to target at least one of the hardware abstraction layer (1180) and / or the application layer (1140) (e.g., a position tracker (1171), a space recognizer (1172), a gesture tracker (1173), an eye-gaze tracker (1174), and / or a face tracker (1175)) may be included. The programs included in the framework layer (1150) may provide an application programming interface (API) that is executable (or callable) based on other programs.

[0157] For example, a program designed to target users of the electronic device (101) may be included within the application layer (1140). As an example of programs included in the application layer (1140), an extended reality (XR) system user interface (UI) (1141) or an XR application (1142) is exemplified, but the embodiments of the present disclosure are not limited thereto. For example, programs (e.g., software applications) included in the application layer (1140) may call an API to cause execution of functions supported by programs included in the framework layer (1150).

[0158] For example, the electronic device (101) may display one or more visual objects on the display (230) for performing interaction with the user based on the execution of the XR system UI (1141). A visual object may refer to an object that can be placed within a screen for transmitting and / or interacting with information, such as text, an image, an icon, a video, a button, a checkbox, a radio button, a text box, a slider, and / or a table. A visual object may be referred to as a visual guide, a virtual object, a visual element, a UI element, a view object, and / or a view element. The electronic device (101) may provide the user with functions available within a virtual space based on the execution of the XR system UI (1141).

[0159] Although FIG. 11 illustrates including a lightweight renderer (1143) and / or an XR plug-in (1144) within the XR system UI (1141), the embodiments of the present disclosure are not limited thereto. For example, based on the XR system UI (1141), the processor (1110) may execute a lightweight renderer (1143) and / or an XR plug-in (1144) within the framework layer (1150).

[0160] For example, the electronic device (101) may obtain resources (e.g., APIs, system processes, and / or libraries) used to define, create, and / or execute a rendering pipeline that allows partial changes based on the execution of a lightweight renderer (1143). The lightweight renderer (1143) may be referred to as a lightweight render pipeline from the perspective of defining a rendering pipeline that allows partial changes. The lightweight renderer (1143) may include a renderer built prior to the execution of a software application (e.g., a prebuilt renderer). For example, the electronic device (101) may obtain resources (e.g., APIs, system processes, and / or libraries) used to define, create, and / or execute an entire rendering pipeline based on the execution of an XR plug-in (1144). The XR plug-in (1144) may be referred to as an open XR native client from the perspective of defining (or configuring) an entire rendering pipeline.

[0161] For example, the electronic device (101) may display a screen representing at least a portion of a virtual space on the display (230) based on the execution of the XR application (1142). The XR plug-in (1144-1) included in the XR application (1142) may include instructions that support functions similar to those of the XR plug-in (1144) of the XR system UI (1141). Descriptions of the XR plug-in (1144-1) that overlap with the description of the XR plug-in (1144) may be omitted. The electronic device (101) may cause the execution of the virtual space manager (1151) based on the execution of the XR application (1142).

[0162] For example, the electronic device (101) may display an image on the display (230) in a virtual space based on the execution of the application (1145). The application (1145) may be configured to output image information for displaying a two-dimensional image. The electronic device (101) may cause the execution of the virtual space manager (1151) based on the execution of the application (1145). The electronic device (101) may generate dual image information to display the two-dimensional image in a three-dimensional virtual space based on the execution of the application (1145). Here, the dual image information may include first image information for the left eye and second image information for the right eye, taking into account binocular parallax. In order to display the two-dimensional image in the three-dimensional virtual space, the electronic device (101) may generate the dual image information based on the image information for displaying the two-dimensional image.

[0163] According to one embodiment, the electronic device (101) may provide a virtual space service based on the execution of the virtual space manager (1151). For example, the virtual space manager (1151) may include a platform for supporting the virtual space service. Based on the execution of the virtual space manager (1151), the electronic device (101) may identify a virtual space formed based on the user's location indicated by data acquired through the sensor (1130), and may display at least a portion of the virtual space on the display (230). The virtual space manager (1151) may be referred to as a composition presentation manager (CPM).

[0164] For example, the virtual space manager (1151) may include a runtime service (1152). As an example, the runtime service (1152) may be referred to as an OpenXR runtime module (or an OpenXR runtime program). The electronic device (101) may execute at least one of a user's pose prediction function, a frame timing function, and / or a spatial input function based on the execution of the runtime service (1152). As an example, the electronic device (101) may perform rendering for a virtual space service to the user based on the execution of the runtime service (1152). For example, a function related to a virtual space, executable by the application layer (1140), may be supported based on the execution of the runtime service (1152).

[0165] For example, the virtual space manager (1151) may include a pass-through manager (1153). Based on the execution of the pass-through manager (1153), the electronic device (101) may display an image and / or video representing an actual space acquired through an external camera on at least a portion of the screen while displaying a screen representing a virtual space (e.g., the screen of FIG. 1A) on the display (230).

[0166] For example, the virtual space manager (1151) may include an input manager (1154). The electronic device (101) may identify data (e.g., sensor data) obtained by executing one or more programs included in the recognition service layer (1170) based on the execution of the input manager (1154). The electronic device (101) may use the obtained data to identify a user input related to the electronic device (101). The user input may be related to a motion (e.g., a hand gesture), gaze, and / or speech of the user identified by a sensor (1120) (e.g., an image sensor (1130) such as an external camera). The user input may be identified based on an external electronic device connected (or paired) via a communication circuit.

[0167] For example, the perception abstract layer (1160) can be used for data exchange between the virtual space manager (1151) and the perception service layer (1170). From the perspective of being used for data exchange between the virtual space manager (1151) and the perception service layer (1170), the perception abstract layer (1160) can be referred to as an interface. For example, the perception abstract layer (1160) can be referenced as OpenPX. The perception abstract layer (1160) can be used for a perception client and a perception service.

[0168] According to one embodiment, the recognition service layer (1170) may include one or more programs for processing data acquired from the sensor (1120). The one or more programs may include at least one of a position tracker (1171), a space recognizer (1172), a gesture tracker (1173), and / or an eye tracker (1174). The type and / or number of the one or more programs included in the recognition service layer (1170) are not limited to those illustrated in FIG. 11.

[0169] For example, the electronic device (101) can identify the pose of the electronic device (101) using the sensor (1130) based on the execution of the position tracker (1171). The electronic device (101) can identify the 6 degrees of freedom pose (6 dof pose) of the electronic device (101) using data acquired using an external camera (e.g., an image sensor (1121)) and / or an IMU (e.g., a motion sensor (1122) including at least one of a gyro sensor, an acceleration sensor, and / or a geomagnetic sensor) based on the execution of the position tracker (1171). The position tracker (1171) may be referred to as a head tracking (HeT) module (or head tracker, head tracking program).

[0170] For example, the electronic device (101) may obtain information for providing a three-dimensional virtual space corresponding to the surrounding environment (e.g., external space) of the electronic device (101) (or the user of the electronic device (101)) based on the execution of the space recognizer (1172). The electronic device (101) may reproduce the surrounding environment of the electronic device (101) in three dimensions using data obtained using an external camera (e.g., an image sensor (1121)) based on the execution of the space recognizer (1172). The electronic device (101) may identify at least one of a plane, a slope, and stairs based on the surrounding environment of the electronic device (101) reproduced in three dimensions based on the execution of the space recognizer (1172). The space recognizer (1172) may be referred to as a scene understanding (SU) module (or a scene recognition program).

[0171] For example, the electronic device (101) can identify (or recognize) a pose and / or gesture of a hand of a user of the electronic device (101) based on the execution of the gesture tracker (1173). As an example, the electronic device (101) can identify a pose and / or gesture of a hand of a user using data acquired from an external camera (e.g., an image sensor (1121)) based on the execution of the gesture tracker (1173). As an example, the electronic device (101) can identify a pose and / or gesture of a hand of a user based on data (or images) acquired using an external camera based on the execution of the gesture tracker (1173). The gesture tracker (1173) may be referred to as a hand tracking (HaT) module (or hand tracking program) and / or a gesture tracking module.

[0172] For example, the electronic device (101) can identify (or track) eye movements of a user of the electronic device (101) based on the execution of the gaze tracker (1174). As an example, the electronic device (101) can identify eye movements of the user using data acquired from a gaze tracking camera (e.g., an image sensor (1121)) based on the execution of the gaze tracker (1174). The gaze tracker (1174) may be referred to as an eye tracking (ET) module (or eye tracking program) and / or a gaze tracking module.

[0173] For example, the recognition service layer (1170) of the electronic device (101) may further include a face tracker (1175) for tracking the user's face. For example, the electronic device (101) may identify (or track) the movement of the user's face and / or the user's expression based on the execution of the face tracker (1175). The electronic device (101) may estimate the user's expression based on the movement of the user's face based on the execution of the face tracker (1175). As an example, the electronic device (101) may identify the movement of the user's face and / or the user's expression based on data (e.g., images and / or videos) acquired using a camera (1125) (e.g., a camera facing at least a portion of the user's face) based on the execution of the face tracker (1175).

[0174] Referring to FIG. 11, the renderer (1190) may include instructions for rendering images in a three-dimensional virtual space. The processor (1110) executing the renderer (1190) may obtain at least one image to be at least partially displayed in the display area of ​​the display (230) in a software application. For example, the processor (1110) executing the renderer (1190) may determine the location of the area in which an application (e.g., XR application (1142), application (1145)) is to be rendered. The processor (1110) executing the renderer (1190) may generate an image of the application to be displayed on the display (230). The renderer (1190) may synthesize images to generate a composite image to be displayed on the display (230).

[0175] For example, the processor (1110) executing the renderer (1190) can divide the display area of ​​the display (230) into a foveated portion (or may be referred to as the foveated area) and a peripheral portion (or may be referred to as the residual area) using the gaze position calculated using the position tracker (1171) and / or the gaze tracker (1174). For example, the processor (1110) detecting the coordinate values ​​of the gaze position can determine the portion of the display area including the coordinate values ​​as the foveated area. The processor (1110) executing the renderer (1190) can obtain at least one image corresponding to each of the foveated area and the residual area, and having a size smaller than the size of the entire display area of ​​the display (230) or a resolution smaller than the resolution of the display area.

[0176] The processor (1110) executing the renderer (1190) may obtain or generate a composite image to be displayed on the display (230) by synthesizing an image corresponding to the foveated area and an image corresponding to the peripheral area. For example, the processor (1110) may perform upscaling to enlarge the image corresponding to the peripheral area to the size of the entire display area of ​​the display (230). On the enlarged image, the processor (1110) may combine the image corresponding to the foveated area to generate a composite image to be displayed on the display (230). Along the boundary line of the image corresponding to the foveated area, the processor (1110) may apply a visual effect, such as blur, to blend the enlarged image and the image corresponding to the foveated area.

[0177] Fig. 12 shows an example of a block diagram of an electronic device (e.g., the electronic device (101) of Figs. 1A to 11) for displaying an image in a virtual space. Fig. 12 describes an example in which a plurality of programs / instructions for displaying an image in a virtual space are executed. The plurality of programs / instructions may all be executed in one processor (e.g., an AP) or may be executed by a plurality of processors (e.g., an AP, a GPU (graphics processing unit), an NPU (neural processing unit)). The meaning of being executed by the plurality of processors means that some programs / instructions may be executed by a first processor and other some programs / instructions may be executed by a second processor different from the first processor.

[0178] Referring to FIG. 12, the electronic device (101) may execute a virtual space manager (1250) (e.g., the virtual space manager (1151) of FIG. 11, CPM) to render an image in a virtual space. For the virtual space manager (1250), at least some of the descriptions of the virtual space manager (1151) of FIG. 11 may be referenced. The virtual space manager (1250) may include a platform for supporting a virtual space service. The virtual space manager (1250) may include a runtime service (1251) (e.g., OpenXR Runtime), a panel renderer (1252) (e.g., 2D Panel Render), and an XR compositor (1253). The electronic device (101) may execute at least one of a user's pose prediction function, a frame timing function, and / or a spatial input function based on the execution of the runtime service (1251). For the runtime service (1251), at least some of the descriptions of the runtime service (1152) of FIG. 11 may be referred to. The electronic device (101) may display at least one image (video) on a panel (e.g., a 2D panel) to implement a virtual space through the display based on the execution of the panel rendering (1252). For example, the electronic device (101) may display a rendering image corresponding to RGB information (1266) for the panel from the spatialization manager (1240) described below through the display (e.g., the display (230)).

[0179] In one embodiment, the electronic device (101) may synthesize an image of an actual area captured by a camera in a virtual space (hereinafter, referred to as a pass-through image) with a virtual area image based on the execution of the XR compositor (1253). For example, the electronic device (101) may generate a composite image by merging the pass-through image and the virtual area image based on the execution of the XR compositor (1253). The electronic device (101) may transmit the generated composite image to a display buffer so that the composite image is displayed. The electronic device (101) may identify a virtual space through the virtual space manager (1250) and display at least a portion of the virtual space on the display (230). The virtual space manager (1250) may be referred to as a CPM. The electronic device (101) may execute the virtual space manager (1250) to render an image corresponding to at least a portion of the virtual space.

[0180] According to one embodiment, the electronic device (101) may execute a spatialization manager (1240). The spatialization manager (1240) may perform processes for displaying an image in a three-dimensional virtual space. The electronic device (101) may perform preprocessing based on the execution of the spatialization manager (1240) so that the image can be rendered in a three-dimensional virtual space through the virtual space manager (1250). For example, the electronic device (101) may perform at least some of the functions of the renderer (1190) of FIG. 11 based on the execution of the spatialization manager (1240). The electronic device (101) may process image information provided by an application (e.g., an XR application (1210), an application (1220) that provides a general 2D screen other than XR, and an application that provides a system UI (1230)) based on the execution of the spatialization manager (1240). A spatialization manager (1240) (e.g., Space Flinger) may include a system scene manager (1241) (e.g., System scene), an input manager (1242) (e.g., Input Routing), and a lightweight rendering engine (1243) (e.g., Impress Engine). The system scene manager (1241) may be executed to display a system UI (1230). System UI-related information (1264) may be transmitted to the system scene manager (1241) from a program (e.g., API) that provides the system UI (1230). The system UI-related information (1264) may be obtained through a spatializer API and / or a same-process private API. The spatialization manager (1240) may determine the layout (e.g., location, display order) of the screen of the system UI (1230) in a 10-dimensional space through pre-allocated resources.The system screen manager (1241) may transmit image information (1267) for rendering the screen of the system UI (1230) to the virtual space manager (1250) according to the layout. The input manager (1242) may be configured to process user input (e.g., user input on a system screen or an app screen). The impression engine (1243) may be a renderer for image generation (e.g., a lightweight renderer (1143)). For example, the impression engine (1243) may be used to display the system UI (1230). According to one embodiment, the spatialization manager (1240) may include a lightweight rendering engine (1243) for rendering the system UI. According to one embodiment, when the lightweight rendering engine (1243) does not have sufficient resources to render an avatar used in the HMD, at least one external rendering engine may be used. At this time, to resolve compatibility issues with external rendering (e.g., 3rd party engines), an external rendering engine support module may be added within the spatialization manager (1240).

[0181] According to one embodiment, the electronic device can execute an application. For example, in response to the execution of an XR application (1210) (e.g., an XR application (1142), a 3D game, an XR map, or other immersive application), the electronic device can execute a virtual space manager (1250). The electronic device (101) can provide dual image information (1261) provided from the XR application (1210) to the virtual space manager (1250). In order to display an image in a three-dimensional space, the dual image information (1261) can include two pieces of image information that take binocular parallax into account. For example, the dual image information (1261) can include first image information for the user's left eye and second image information for the user's right eye for rendering in a three-dimensional virtual space. Hereinafter, in the present disclosure, the term dual image information is used to refer to image information for displaying images for both eyes in a three-dimensional space. In addition to dual image information, the above dual image information may also include binocular image information, dual image information, dual image data, dual images, binocular image data, stereoscopic image information, 3D image information, spatial image information, spatial image data, 2D-3D conversion data, dimensional conversion image data, binocular parallax image data, and / or equivalent technical terms. The electronic device (101) can generate a composite image by merging image layers through a virtual space manager (1250). The electronic device (101) can transmit the generated composite image to a display buffer. The composite image can be displayed on the display (230) of the electronic device (101).

[0182] According to one embodiment, the electronic device can execute at least one application among an XR application (1210) and other applications (1220) (e.g., a first application (1220-1), a second application (1220-2), ..., an Nth application (1220-N)). According to one embodiment, the application (1220) can be configured to output image information for displaying a two-dimensional image. In other words, the application (1220) can provide a two-dimensional image. For example, the application (1220) can be a video application, a schedule application, or an application (1220) can be an Internet browser application. If, in response to the execution of the application (1220), the image information (1262) provided from the application (1220) is provided to the virtual space manager (1250). Since the image information (1262) has only the x-coordinate and y-coordinate within a two-dimensional plane, it may be difficult to consider the chronological relationship (i.e., the distance from the user) between other applications centered on the user. Even when displaying an application (1220) that provides a general 2D screen, the electronic device (101) may execute the spatialization manager (1240) to provide dual image information to the virtual space manager (1250). For example, based on the execution of the spatialization manager (1240), the electronic device (101) may receive application-related information (1263) from the first application (1220-1). For example, the application-related information (1263) may include image information representing a two-dimensional image of the first application (1220-1) (e.g., information including RGB for each pixel) and / or content information in the first application (1220-1) (e.g., characteristics of content executed in the first application, type of content). Application related information (1263) can be obtained through the spatializer API.Based on the execution of the spatialization manager (1240), the electronic device (101) can identify information about the location of the area to be rendered by the first application (1220-1) and the size of the area to be rendered (hereinafter, location information). Based on the execution of the spatialization manager (1240), the electronic device (101) can generate dual image information (1265, e.g., RGBx2) that takes into account the user's binocular disparity through the image information and the location information. Based on the execution of the spatialization manager (1240), the electronic device (101) can provide the dual image information (1265) to the virtual space manager (1250). By converting a simple two-dimensional image into the dual image information (1265), a problem that occurs when the image information (1262) is directly transmitted to the virtual space manager (1250) can be resolved. Additionally, since at least some of the functions for displaying images in a virtual space are performed by the spatialization manager (1240) instead of the virtual space manager (1250), the burden on the virtual space manager (1250) can be reduced.

[0183] In one embodiment, a method for storing depth values ​​for stereoscopic rendering of a two-dimensional image using an alpha channel may be required. According to one embodiment, an electronic device as described above may include a memory (and / or including one or more storage media) for storing instructions, and at least one processor including a processing circuit. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain depth information of a visual object. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify depth values ​​to be included in an alpha channel representing transparency of the visual object using the depth information. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate an image representing the visual object, wherein the depth values ​​and transparencies are each included in an alpha channel of pixels.

[0184] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify a range of the depth values. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to determine, according to the range, a number of bits to be occupied to represent the depth values ​​within the alpha channel and a number of bits to be occupied to represent the transparencies. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate, according to the determination, the image including the depth values ​​and the transparencies within the alpha channel.

[0185] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate the image, using the depth information, including the depth values ​​represented by the first number of bits and the transparencies represented by a second number of bits obtained by subtracting the first number of bits from a total number of bits included in the alpha channel, based on determining that the depth values ​​are represented by a first number of bits. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate the image, using the depth information, including the depth values ​​represented by the third number of bits and the transparencies represented by a fourth number of bits obtained by subtracting the third number of bits from the total number, based on determining that the depth values ​​are represented by a third number of bits greater than the first number of bits.

[0186] For example, within the alpha channel of each of the pixels, a bit sequence representing a depth value may be located behind the least significant bit (LSB) of a bit sequence representing transparency within the alpha channel.

[0187] For example, within the alpha channel of each of the pixels, a bit sequence representing a depth value may be positioned before the most significant bit (MSB) of a bit sequence representing transparency within the alpha channel.

[0188] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate metadata indicating a number of bits within the alpha channel reserved for representing each of the depth values. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate a file including the metadata and the image.

[0189] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to add the depth values ​​and the transparencies to the alpha channel of pixels corresponding to the first region while generating the image, the image including a first region corresponding to the visual object and a second region surrounding the first region. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to add, to pixels corresponding to the second region, bit numbers of the depth values ​​added to the alpha channel of the pixels corresponding to the first region.

[0190] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain different depth information based on a shape of the visual object at a second point in time subsequent to the first point in time corresponding to the first image, based on generating a video including the image as a first image. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate a second image, the second image including, in an alpha channel of pixels, differences between depth values ​​represented by the different depth information and depth values ​​included in the alpha channel of the first image, respectively. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate the video including the first image and the second image.

[0191] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate the at least one image, wherein difference values ​​for the depth values ​​included in the alpha channel of the first image are each included in the alpha channel of pixels, based on generating the first image corresponding to a key frame within the video, and generating at least one image corresponding to a time interval of a specified length associated with the key frame from the first time point corresponding to the first image.

[0192] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate a video including the image as a first image, based on a specified number of images rendered after the first image corresponding to a key frame within the video, wherein difference values ​​for the depth values ​​included in the alpha channel of the first image are each included in the alpha channel of pixels.

[0193] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain other depth information associated with the visual object at a second point in time subsequent to a first point in time corresponding to the first image, based on generating a video including the image as a first image. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain difference values ​​between depth values ​​included in an alpha channel of the pixels of the first image, represented by the depth information, and other depth values ​​corresponding to each of pixels of a second image corresponding to the second point in time, represented by the other depth information. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate the second image, wherein the difference values ​​are each included in the alpha channel of the pixels, based on obtaining the difference values ​​included in a reference range. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate the second image, wherein the different depth values ​​and transparencies are each included in the alpha channel of the pixels, based on obtaining the difference values ​​outside the reference range.

[0194] For example, the electronic device may include a sensor configured to detect a motion of a user. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain, using sensor data detected from the sensor at a first time point, the depth information used to generate the first image based on representing the visual object moving in accordance with the motion detected by the sensor and generating a video including the image as a first image. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to detect sensor data from the sensor at a second time point after the first time point. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify a difference between the sensor data detected at the first time point and the sensor data detected at the second time point. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate a second image corresponding to the second viewpoint, wherein the difference values ​​of the depth values ​​included in the alpha channel of the first image and other depth values ​​represented by other depth information acquired based on the sensor data of the second viewpoint are respectively included in the alpha channel of pixels, based on identifying the difference included in the reference range. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate the second image corresponding to the second viewpoint, wherein the other depth values ​​are respectively included in the alpha channel of the pixels, based on identifying the difference outside the reference range.

[0195] For example, the electronic device may include a display assembly including a plurality of displays. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive an input for displaying the image. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain, based on the input, the depth values ​​included in the alpha channel of the pixels of the image. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to determine, based on the obtained depth values, a binocular parallax of each of the pixels. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display, based on the binocular parallax, the image on a first display of the plurality of displays. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display, on a second display of the plurality of displays, another image representing the visual object shifted based on the binocular disparity.

[0196] For example, the visual object may include an avatar representing a user of the electronic device.

[0197] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to begin generating the image using a virtual space including the avatar based on receiving input for rendering the avatar.

[0198] As described above, in one embodiment, a method of an electronic device may be provided. The method may include an operation of obtaining depth information of a visual object. The method may include an operation of identifying depth values ​​to be included in an alpha channel representing transparency of the visual object using the depth information. The method may include an operation of generating an image representing the visual object, wherein the depth values ​​and transparencies are each included in an alpha channel of pixels.

[0199] For example, the identifying operation may include an operation of identifying a range of the depth values. The method may include an operation of determining, based on the range, a number of bits to be occupied for representing the depth values ​​within the alpha channel and a number of bits to be occupied for representing the transparencies. The generating operation may include an operation of generating, based on the determination, an image within the alpha channel, the image including the depth values ​​and the transparencies.

[0200] For example, the generating operation may include generating the image, including the depth values ​​expressed in bits of the first number of bits, using the depth information, based on determining that the depth values ​​are represented in bits of the first number of bits, and the transparencies expressed in bits of a second number of bits obtained by subtracting the first number of bits from the total number of bits included in the alpha channel. The method may include generating the image, including the depth values ​​expressed in bits of the third number of bits, using the depth information, based on determining that the depth values ​​are represented in bits of a third number of bits greater than the first number of bits, and the transparencies expressed in bits of a fourth number of bits obtained by subtracting the third number of bits from the total number.

[0201] For example, within the alpha channel of each of the pixels, a bit sequence representing a depth value may be located behind the least significant bit (LSB) of a bit sequence representing transparency within the alpha channel.

[0202] For example, within the alpha channel of each of the pixels, a bit sequence representing a depth value may be positioned before the most significant bit (MSB) of a bit sequence representing transparency within the alpha channel.

[0203] For example, the generating operation may include generating metadata indicating the number of bits within the alpha channel reserved for representing each of the depth values. The method may include generating a file including the metadata and the image.

[0204] For example, the generating operation may include adding the depth values ​​and the transparencies to the alpha channel of pixels corresponding to the first area while generating the image, which includes a first area corresponding to the visual object and a second area surrounding the first area. The method may include adding, to pixels corresponding to the second area, bit numbers of the depth values ​​added to the alpha channel of the pixels corresponding to the first area.

[0205] For example, the method may include an operation of generating a video including the image, which is a first image. The operation of generating the video may include an operation of obtaining different depth information based on a shape of the visual object at a second time point after a first time point corresponding to the first image. The method may include an operation of generating a second image, in which difference values ​​between depth values ​​represented by the different depth information and depth values ​​included in the alpha channel of the first image are each included in an alpha channel of pixels. The method may include an operation of generating the video including the first image and the second image.

[0206] For example, the act of generating the video may include generating the first image corresponding to a key frame within the video, and generating at least one image corresponding to a time interval of a specified length related to the key frame from the first time point corresponding to the first image, wherein difference values ​​for the depth values ​​included in the alpha channel of the first image are each included in the alpha channel of pixels.

[0207] For example, the method may include generating a video including the image, which is a first image. The generating of the video may include generating a specified number of images, each of which includes depth values ​​included in the alpha channel of pixels, based on a specified number of images rendered after the first image corresponding to a key frame within the video.

[0208] For example, the method may include generating a video including the image, which is a first image. The generating the video may include obtaining other depth information associated with the visual object at a second time point after a first time point corresponding to the first image. The method may include obtaining difference values ​​between depth values ​​included in an alpha channel of the pixels of the first image, which are represented by the depth information, and other depth values ​​corresponding to each of the pixels of the second image, which are represented by the other depth information, corresponding to the second time point. The method may include generating the second image, based on obtaining the difference values ​​included in a reference range, in which the difference values ​​are respectively included in the alpha channel of the pixels. The method may include generating the second image, based on obtaining the difference values ​​outside the reference range, in which the other depth values ​​and transparencies are respectively included in the alpha channel of the pixels.

[0209] For example, the method may include an operation of representing the visual object moving according to the motion detected by a sensor configured to detect a motion of a user, and generating a video including the image as a first image. The operation of generating the video may include an operation of obtaining the depth information used for generating the first image using sensor data detected from the sensor at a first time point. The method may include an operation of detecting sensor data from the sensor at a second time point after the first time point. The method may include an operation of identifying a difference between the sensor data detected at the first time point and the sensor data detected at the second time point. The method may include an operation of generating a second image corresponding to the second time point, wherein difference values ​​of other depth values ​​represented by the depth values ​​included in the alpha channel of the first image and other depth information acquired based on the sensor data at the second time point are each included in an alpha channel of pixels based on identifying the difference included in a reference range. The method may include generating the second image corresponding to the second viewpoint, wherein the different depth values ​​are each included in the alpha channel of the pixels, based on identifying the difference outside the reference range.

[0210] For example, the method, when performed by the electronic device further comprising a display assembly including a plurality of displays, may include receiving an input for displaying the image. The method may include obtaining, based on the input, the depth values ​​included in the alpha channels of the pixels of the image. The method may include determining, based on the obtained depth values, a binocular parallax of each of the pixels. The method may include displaying, based on the binocular parallax, the image on a first display among the plurality of displays. The method may further include displaying, on a second display among the plurality of displays, another image representing the visual object shifted based on the binocular parallax.

[0211] For example, the visual object may include an avatar representing a user of the electronic device.

[0212] For example, the action of generating the image may be initiated using a virtual space containing the avatar in response to receiving an input for rendering the avatar.

[0213] In one embodiment, as described above, a non-transitory computer-readable storage medium storing instructions may be provided. The instructions, when executed by an electronic device including a display assembly including a plurality of displays, may cause the electronic device to obtain a file including an image representing a visual object. The instructions, when executed by the electronic device, may cause the electronic device to identify, from an alpha channel of pixels of the image, depth values ​​and transparencies of portions of the visual object corresponding to each of the pixels. The instructions, when executed by the electronic device, may cause the electronic device to control the plurality of displays such that, while controlling the display assembly to display the visual object represented based on the transparencies, the portions displayed on a first display of the plurality of displays are respectively shifted from the portions displayed on a second display of the plurality of displays according to the depth values.

[0214] According to one embodiment, an electronic device as described above may include a display assembly including a plurality of displays, a memory storing instructions and including one or more storage media, and at least one processor including a processing circuit. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain a file including an image representing a visual object. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify, from an alpha channel of pixels of the image, depth values ​​and transparencies of portions of the visual object corresponding to each of the pixels. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to control the plurality of displays such that, while controlling the display assembly, the portions displayed on a first display of the plurality of displays are shifted from the portions displayed on a second display of the plurality of displays, respectively, according to the depth values, to display the visual object expressed based on the transparencies.

[0215] According to one aspect of the present disclosure, an electronic device includes a memory including one or more storage media storing instructions, and at least one processor including a processing circuit, wherein the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to obtain depth information of a visual object, identify depth values ​​based on the depth information of the visual object, add the depth values ​​to an alpha channel of an image representing the visual object, and generate the image, wherein the alpha channel includes the depth values ​​and transparencies of the visual object.

[0216] The instructions, when individually or collectively executed by the at least one processor, may further cause the electronic device to identify a range of depth values, determine a first number of bits representing the depth values ​​and a second number of bits representing the transparencies within the alpha channel based on the range of depth values, and generate the image based on the determined first number of bits and the determined second number of bits.

[0217] In a first case where the depth values ​​are represented by first bits of the first number of bits, and the first bits are determined using the depth information, the transparencies can be represented by second bits of the second number of bits, and the second number of bits is a total number of bits included in the alpha channel, by subtracting the first bits of the first number of bits, and in a second case where the depth values ​​are represented by bits of a third number of bits that is greater than the first number of bits, and the third bits are determined using the depth information, the transparencies can be represented by fourth bits of a fourth number of bits, and the fourth number of bits is a total number of bits included in the alpha channel, by subtracting the third bits of the third number of bits.

[0218] Within the alpha channel, a first bit sequence representing the depth values ​​may be positioned behind a least significant bit (LSB) of a second bit sequence representing the transparencies within the alpha channel.

[0219] A first bit sequence representing the depth values ​​within the alpha channel may be positioned before a most significant bit (MSB) of a second bit sequence representing the transparency within the alpha channel.

[0220] The instructions, when individually or collectively executed by the at least one processor, may further cause the electronic device to generate metadata indicating a number of bits in the alpha channel, the metadata being reserved to indicate depth values, and to generate a file including the metadata and the image.

[0221] The image may include a first region corresponding to the visual object, and a second region surrounding the first region, and the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to add the depth values ​​and the transparencies to first pixels of the alpha channel, and to add the number of bits of the depth values ​​added to the first pixels to second pixels of the alpha channel, the first pixels corresponding to the first region and the second pixels corresponding to the second region.

[0222] The image may be a first image, and the instructions, when individually or collectively executed by the at least one processor, may further cause the electronic device to obtain different depth information based on a shape of the visual object at a second point of time after a first point of time corresponding to the first image, generate a second image having different depth values ​​represented by the different depth information, and an alpha channel including difference values ​​between the depth values ​​in the alpha channel of the first image, and generate a video including the first image and the second image.

[0223] The instructions, when individually or collectively executed by the at least one processor, may further cause the electronic device to generate the at least one image, having the alpha channel including difference values ​​for the depth values ​​included in the alpha channel of the first image, based on generating the first image corresponding to a key frame within the video, and based on generating at least one image corresponding to a time interval of a specified length associated with the key frame from the first time point corresponding to the first image.

[0224] The image may be a first image, and the instructions, when individually or collectively executed by the at least one processor, may further cause the electronic device to generate a specified number of images, including an alpha channel including difference values ​​for the depth values ​​within the alpha channel of the first image, based on a specified number of images rendered after the first image corresponding to a key frame within the video.

[0225] The image may be a first image, and the instructions, when individually or collectively executed by the at least one processor, may further cause the electronic device to obtain other depth information related to the visual object at a second point of time subsequent to a first point of time corresponding to the first image, obtain difference values ​​between depth values ​​in the alpha channel of the first image, which are represented by the depth information, and other depth values, which are respectively corresponding to a second image corresponding to the second point of time, which are represented by the other depth information, and generate the second image having the alpha channel including the difference values ​​based on obtaining the difference values ​​included in a reference range, and generate the second image having the alpha channel having the other depth values ​​and the other transparencies based on obtaining the difference values ​​outside the reference range.

[0226] The electronic device may further include a sensor configured to detect a motion of a user, and the image is a first image representing the visual object moved based on the motion detected by the sensor, and the instructions, when individually or collectively executed by the at least one processor, may further cause the electronic device to obtain the depth information to generate the first image using first sensor data detected from the sensor at a first time point, detect second sensor data from the sensor at a second time point after the first time point, identify a difference between the first sensor data detected at the first time point and the second sensor data detected at the second time point, and generate a second image corresponding to the second time point based on identifying the difference included in a reference range, and the alpha channel of the second image includes difference values ​​of other depth values ​​represented by the depth values ​​included in the alpha channel of the first image and other depth information acquired based on the sensor data at the second time point, and generate the second image corresponding to the second time point based on identifying the difference outside the reference range. and the alpha channel of the second image contains the different depth values.

[0227] The electronic device may further include a display assembly including a plurality of displays, and the instructions, when individually or collectively executed by the at least one processor, may further cause the electronic device to receive an input for displaying the image, obtain depth values ​​included in the alpha channel of the image based on receiving the input, determine binocular parallax of the alpha channel based on the obtained depth values, display the image on a first display among the plurality of displays based on the binocular parallax, and display another image representing the visual object shifted based on the binocular parallax on a second display among the plurality of displays based on the binocular parallax.

[0228] The visual object may include an avatar representing a user of the electronic device.

[0229] The instructions, when individually or collectively executed by the at least one processor, may further cause the electronic device to initiate generation of the image using a virtual space including the avatar based on receiving an input for rendering the avatar.

[0230] According to one aspect of the present disclosure, a method of an electronic device includes an operation of obtaining depth information of a visual object, an operation of identifying depth values ​​using the depth information, an operation of adding the depth values ​​to an alpha channel of an image representing the visual object, and an operation of generating the image, wherein the alpha channel includes the depth values ​​and transparencies of the visual object.

[0231] Based on the depth information of the visual object, the operation of identifying the depth values ​​may include an operation of identifying a range of the depth values, an operation of determining a first number of bits representing the depth values ​​within the alpha channel, and a second number of bits representing the transparencies within the alpha channel, and the operation of generating the image may include an operation of generating the image based on the determined first number of bits and the determined second number of bits.

[0232] In a first case, wherein the depth values ​​are represented by first bits of the first number of bits, and the first bits are determined using the depth information, the transparencies can be represented by second bits of the second number of bits, and the second number of bits is determined such that the first bits of the first number of bits are subtracted from the total number of bits included in the alpha channel, and the depth values ​​are represented by third bits of a third number of bits that is greater than the first number of bits, and the third bits are determined using the depth information, in a second case, wherein the transparencies can be represented by fourth bits of the fourth number of bits, and the fourth number of bits is determined such that the third bits of the third number of bits are subtracted from the total number of bits of the alpha channel.

[0233] Within the alpha channel, a first bit sequence representing the depth values ​​may be positioned behind a least significant bit (LSB) of a second bit sequence representing the transparencies within the alpha channel.

[0234] Within the alpha channel, a first bit sequence representing depth values ​​may be positioned before a most significant bit (MSB) of a second bit sequence representing transparencies within the alpha channel.

[0235] As used herein, the term "if" will be understood to mean "when, upon," "in response to determining," or "in response to detecting," depending on the context. Similarly, "if it is determined to," or "if [the stated condition or event] is detected," will optionally be understood to mean "upon determining," or "in response to determining," "upon detecting [the stated condition or event]," or "in response to detecting [the stated condition or event]."

[0236] The devices described above may be implemented as hardware components, software components, and / or a combination of hardware components and software components. For example, the devices and components described in the embodiments may be implemented using one or more general-purpose computers or special-purpose computers, such as a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing instructions and responding to them. The processing device may execute an operating system (OS) and one or more software applications running on the operating system. The processing device may also access, store, manipulate, process, and generate data in response to the execution of the software. While the processing device is sometimes described as being used singly, those skilled in the art will recognize that the processing device may include multiple processing elements and / or multiple types of processing elements. For example, the processing device may include multiple processors, or one processor and one controller. Additionally, other processing configurations, such as parallel processors, are also possible.

[0237] Software may include a computer program, code, instructions, or a combination of one or more of these, which may configure a processing device to perform a desired operation or may independently or collectively command the processing device. The software and / or data may be embodied in any type of machine, component, physical device, computer storage medium, or device for interpretation by the processing device or for providing instructions or data to the processing device. The software may also be distributed over networked computer systems and stored or executed in a distributed manner. The software and data may be stored on one or more computer-readable recording media.

[0238] According to one embodiment, the method may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. In this case, the medium may be one that continuously stores a computer-executable program or one that temporarily stores it for execution or download. In addition, the medium may be various recording or storage means in the form of a single or multiple hardware combinations, and is not limited to a medium directly connected to a computer system, but may also be distributed over a network. Examples of the medium may include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and those configured to store program instructions, including ROM, RAM, and flash memory. In addition, examples of other media may include recording or storage media managed by app stores that distribute applications, sites that supply or distribute various software, servers, etc.

[0239] Although the embodiments described above have been described by way of limited examples and drawings, those skilled in the art will appreciate that various modifications and variations can be made based on the above teachings. For example, appropriate results can still be achieved even if the described techniques are performed in a different order than described, and / or components of the described systems, structures, devices, circuits, etc. are combined or combined in a different manner than described, or are replaced or substituted with other components or equivalents.

[0240] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims described below.

Claims

1. In electronic devices, A memory including one or more storage media for storing instructions; and At least one processor comprising a processing circuit, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Obtain depth information of visual objects; Based on the depth information of the visual object, depth values ​​are identified; Adding the above depth values ​​to the alpha channel of the image representing the visual object; and causing the above image to be generated, and The alpha channel contains the depth values ​​and the transparency of the visual object. Electronic devices.

2. In claim 1, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Identify the range of the above depth values; Based on the range of the depth values, determining a first number of bits representing the depth values ​​and a second number of bits representing the transparencies within the alpha channel; and Further causing the image to be generated based on the determined first bit number and the determined second bit number, Electronic devices.

3. In claim 2, In a first case where the depth values ​​are represented by first bits of the first number of bits, and the first bits are determined using the depth information, the transparencies are represented by second bits of the second number of bits, and the second number of bits is a total number of bits included in the alpha channel, wherein the first bits of the first number of bits are subtracted; and In the second case, where the depth values ​​are represented by a third number of bits greater than the first number of bits, and the third bits are determined using the depth information, the transparencies are represented by fourth bits of a fourth number of bits, and the fourth number of bits is obtained by subtracting third bits of the third number of bits from the total number of the alpha channel. Electronic devices.

4. In claim 1, within the alpha channel, the first bit sequence representing the depth values ​​is: Within the above alpha channel, located after the LSB (least significant bit) of the second bit sequence representing the above transparencies, Electronic devices.

5. In claim 1, the first bit sequence representing the depth values ​​within the alpha channel is Within the above alpha channel, located before the MSB (most significant bit) of the second bit sequence representing the transparency, Electronic devices.

6. In claim 1, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Generate metadata indicating the number of bits in the alpha channel, wherein the metadata is reserved to indicate depth values; and further causing a file including the metadata and the image to be generated, Electronic devices.

7. In claim 1, The image includes a first area corresponding to the visual object, and a second area surrounding the first area, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Adding the depth values ​​and the transparencies to the first pixels of the alpha channel; and Causes the second pixels of the alpha channel to add the number of bits of the depth values ​​added to the first pixels, The first pixels correspond to the first area and the second pixels correspond to the second area, Electronic devices.

8. In claim 1, the image is a first image, and The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Obtaining different depth information based on the shape of the visual object at a second time point after the first time point corresponding to the first image; Generating a second image having said alpha channel including said other depth values ​​represented by said other depth information and difference values ​​between said depth values ​​in said alpha channel of said first image; and Further causing a video to be generated including the first image and the second image, Electronic devices.

9. In claim 8, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Further causing the method to generate at least one image, having an alpha channel including difference values ​​for the depth values ​​included in the alpha channel of the first image, based on generating the first image corresponding to a key frame within the video, and based on generating at least one image corresponding to a time interval of a specified length related to the key frame from the first time point corresponding to the first image. Electronic devices.

10. In claim 1, the image is a first image, and The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Based on a specified number of images rendered after the first image corresponding to the key frame within the video: Further causing the generation of the specified number of images, wherein the alpha channel includes difference values ​​for the depth values ​​within the alpha channel of the first image. Electronic devices.

11. In claim 1, the image is a first image, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Obtaining other depth information related to the visual object at a second time point after the first time point corresponding to the first image; Obtain difference values ​​between depth values ​​in the alpha channel of the first image, represented by the depth information, and other depth values ​​corresponding to the second image corresponding to the second viewpoint, represented by the other depth information; generating the second image having the alpha channel including the difference values ​​based on obtaining the difference values ​​included in the reference range; and Further causing the second image to be generated based on obtaining the difference values ​​outside the reference range, having the alpha channel with the different depth values ​​and the different transparencies. Electronic devices.

12. In claim 1, further comprising a sensor configured to detect the motion of the user; The above image is a first image representing the visual object that has moved based on the motion detected by the sensor, and The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Obtaining the depth information to generate the first image using the first sensor data detected from the sensor at the first point in time; Detecting second sensor data from the sensor at a second time point after the first time point; Identifying a difference between the first sensor data detected at the first point in time and the second sensor data detected at the second point in time; Based on identifying the difference included in the reference range, generating a second image corresponding to the second time point, wherein the alpha channel of the second image includes difference values ​​of other depth values ​​represented by the depth values ​​included in the alpha channel of the first image and other depth information acquired based on the sensor data of the second time point; and Further causing the second image corresponding to the second time point to be generated based on identifying the difference outside the reference range, wherein the alpha channel of the second image includes the different depth values. Electronic devices.

13. In claim 1, Further comprising a display assembly comprising a plurality of displays, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Receive input for displaying the above image; Based on receiving the above input, obtaining the depth values ​​included in the alpha channel of the image; Based on the acquired depth values, the binocular parallax of the alpha channel is determined; Based on the above binocular disparity, displaying the image on a first display among the plurality of displays; and Further causing, based on the binocular disparity, to display another image, representing the visual object shifted based on the binocular disparity, on a second display among the plurality of displays. Electronic devices.

14. In claim 1, the visual object is: including an avatar representing a user of the electronic device; Electronic devices.

15. In claim 14, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Further causing, based on receiving an input for rendering the avatar, to start generating the image using a virtual space including the avatar. Electronic devices.

Citation Information

Patent Citations

  • Method and apparatus for colorimetric sensor-based urine glucose identification using artificial intelligence technique

    KR1020250067388A

  • Waterproof structure of X-ray detector battery cover

    KR102373970B1

  • Method and Apparatus for Synthesis of Light field using Variable Layered Depth Image

    KR102492163B1

  • Explosion proof apparatus for gas insulated switchgear

    KR102568814B1

  • Method and apparatus for rasterizing and encoding vector graphics

    US9704270B1