Electronic device and method for applying stereoscopic visual effect to image, and non-transitory computer-readable storage medium
By using a processor to execute programs for object detection and depth estimation, the electronic device applies stereoscopic visual effects that enhance depth and immersion by ensuring natural foreground-background interactions, addressing the challenge of effective stereoscopic image rendering.
Patent Information
- Application Number
- PCT/KR2025/003830
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-09
- Filing Date
- 2025-03-25
- Publication Date
- 2026-01-02
AI Technical Summary
Existing electronic devices struggle to effectively apply stereoscopic visual effects to images, particularly in enhancing the sense of depth and immersion by accurately distinguishing foreground and background objects and applying appropriate parallax effects.
The electronic device employs a processor to execute various programs, including a salient object detector, depth information estimator, and 3D renderer, to identify and manipulate image segments, calculate depth values, and apply stereoscopic visual effects based on object positions and characteristics, generating videos with dynamic foreground and background movements.
This approach enhances the sense of depth and immersion by providing natural and immersive stereoscopic visual experiences, ensuring that foreground objects move naturally relative to the background, thus improving user engagement and perception of three-dimensionality.
Smart Images

Figure KR2025003830_02012026_PF_FP_ABST
Abstract
Description
Electronic device, method, and non-transitory computer-readable storage medium for applying stereoscopic visual effects to images
[0001] The present disclosure relates to an electronic device, a method, and a non-transitory computer-readable storage medium for applying a three-dimensional visual effect to an image.
[0002] The shape and / or size of electronic devices are diversifying. To enhance mobility, electronic devices with reduced size and / or volume are being designed. Electronic devices may include cameras for capturing images and / or videos of the external environment. Electronic devices may display images and / or videos captured (or captured) by the cameras.
[0003] In one embodiment, an electronic device may include a display, a memory including one or more storage media for storing instructions, and at least one processor including a processing circuit. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display an image on the display. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive an input for applying a three-dimensional visual effect to the image based on the display of the image. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify segmentation information representing an object in the image based on the input. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify an object from the segmentation information. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate, on the display, a first video representing a background region within the image displaced a first distance beyond the object, based on the object including the edge of the image, thereby applying the stereoscopic visual effect. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate, on the display, a second video representing the background region within the image displaced a second distance from the object, based on the object, thereby applying the stereoscopic visual effect.The second distance may be shorter than the first distance.
[0004] According to one embodiment, a method of an electronic device including a display may include an operation of displaying an image on the display. The method may include an operation of receiving an input for applying a stereoscopic visual effect to the image based on the displaying of the image. The method may include an operation of identifying segmentation information representing an object of the image based on the input. The method may include an operation of identifying the object from the segmentation information. The method may include an operation of generating a first video representing a background area within the image moved a first distance beyond the object on the display based on the object including the edge of the image, and applying the stereoscopic visual effect. The method may include an operation of generating a second video representing the background area within the image moved a first distance beyond the object on the display based on the object spaced from the edge of the image, and applying the stereoscopic visual effect. The second distance may be shorter than the first distance.
[0005] In one embodiment, an electronic device may include a display, a memory including one or more storage media and storing instructions, and at least one processor including a processing circuit. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display an image on the display. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive an input for applying a stereoscopic visual effect to the image based on the displaying of the image on the display. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify segmentation information representing an object in the image based on the input. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate, on the display, a first video representing a background region within the image displaced a first distance beyond the object as a result of applying the stereoscopic visual effect, based on identifying, from the segmentation information, the object including the edge of the image. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate, on the display, a second video representing a background region within the image displaced a second distance beyond the object, which is less than the first distance, as a result of applying the stereoscopic visual effect, based on identifying, from the segmentation information, the object spaced apart from the edge of the image.
[0006] In one embodiment, a method of an electronic device including a display may be provided. The method may include an operation of displaying an image on the display. The method may include an operation of receiving an input for applying a stereoscopic visual effect to the image based on the displaying of the image on the display. The method may include an operation of identifying segmentation information representing an object in the image based on the input. The method may include an operation of generating, on the display, a first video representing a background region within the image displaced by a first distance beyond the object as a result of applying the stereoscopic visual effect, based on identifying, from the segmentation information, the object that includes an edge of the image. The method may include an operation of generating, on the display, a second video representing, on the display, the background region within the image displaced by a second distance less than the first distance beyond the object as a result of applying the stereoscopic visual effect, based on identifying, from the segmentation information, the object that is displaced by the edge of the image.
[0007] In one embodiment, a non-transitory computer-readable storage medium storing instructions may be provided. The instructions, when executed by an electronic device including a display, may cause the electronic device to receive an input for applying a stereoscopic visual effect to an image including an object and a background region. The instructions, when executed by the electronic device, may cause the electronic device to determine, based on the input, one of a plurality of designated stereoscopic visual effects using a position of an object within the image, a depth of the object, and a depth of the background region. The instructions, when executed by the electronic device, may cause the electronic device to apply the determined stereoscopic visual effect to the image, thereby generating a video corresponding to the image. The instructions, when executed by the electronic device, may cause the electronic device to display the generated video on the display.
[0008] In one embodiment, an electronic device may include a display, a memory including one or more storage media for storing instructions, and at least one processor including a processing circuit. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive an input for applying a stereoscopic visual effect to an image including an object and a background region. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to determine, based on the input, one of a plurality of designated stereoscopic visual effects using a position of an object within the image, a depth of the object, and a depth of the background region. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to apply the determined stereoscopic visual effect to the image and generate a video corresponding to the image. The above instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display the generated video on the display.
[0009] Additional aspects will be partly explained by the description which follows, and partly will be obvious from the description, or may be learned by practice of the disclosed embodiments.
[0010] The above-described and other aspects, features, and advantages of some embodiments of the present disclosure will become more apparent from the following description taken in conjunction with the accompanying drawings, in which:
[0011] FIG. 1 illustrates an exemplary operation of an electronic device for applying a three-dimensional visual effect to an image, according to one embodiment;
[0012] FIG. 2 is a block diagram of an electronic device according to one embodiment;
[0013] FIG. 3 illustrates a flow diagram of an electronic device according to one embodiment;
[0014] FIG. 4 illustrates the operation of an electronic device for determining a stereoscopic visual effect to be applied to an image by using the position of the object and / or the depth distribution of the image including the object, according to one embodiment;
[0015] FIG. 5 illustrates the operation of an electronic device for comparing depths of object and background areas, according to one embodiment;
[0016] FIG. 6 illustrates the operation of an electronic device for applying a stereoscopic visual effect to an exemplary image, according to one embodiment;
[0017] FIG. 7 illustrates the operation of an electronic device for applying a stereoscopic visual effect to an exemplary image, according to one embodiment;
[0018] FIG. 8 illustrates the operation of an electronic device for applying a stereoscopic visual effect to an exemplary image, according to one embodiment;
[0019] FIG. 9 illustrates the operation of an electronic device for applying a stereoscopic visual effect to an exemplary image, according to one embodiment;
[0020] FIG. 10 is a flowchart of operations of an electronic device according to one embodiment;
[0021] FIGS. 11A and 11B illustrate exemplary operations of an electronic device to determine a stereoscopic visual effect to be applied to an image, according to one embodiment;
[0022] FIG. 12 illustrates a user interface (UI) displayed by an electronic device to control a stereoscopic visual effect, according to one embodiment;
[0023] FIG. 13 illustrates an exemplary operation of an electronic device for generating depth information from an image, according to one embodiment; and
[0024] FIG. 14 is a block diagram of an electronic device within a network environment according to various embodiments.
[0025] Hereinafter, various embodiments are described with reference to the attached drawings.
[0026] FIG. 1 illustrates an exemplary operation of an electronic device (101) that applies a three-dimensional visual effect to an image (120) according to one embodiment. Referring to FIG. 1, an electronic device (101) including a foldable housing is exemplarily illustrated. The foldable housing (or housing) may include a first housing part (161), a second housing part (162), and a hinge part (163) configured to rotatably couple the first housing part (161) to the second housing part (162). The electronic device (101) may include a display (110) disposed on the first housing part (161) and the second housing part (162). The display (110) may extend from the first housing part (161) across the hinge part (163) to the second housing part (162). The display (110) may be a flexible display. However, the embodiment is not limited thereto, and various form factors of the electronic device (101) are exemplarily described with reference to FIG. 2.
[0027] Referring to an exemplary state (191) of FIG. 1, the electronic device (101) may display an image (120) on the display (110). The image (120) may include a photograph captured by a camera and / or an image sensor. According to one embodiment, the image (120) may be stored in a file having a format based on the Joint Photographic Experts Group (JPEG). The format of the file is not limited to the JPEG, and may have a format of a portable network graphic (PNG), a graphics interchange format (GIF), and / or a Photoshop document (PSD). The file may have a designated extension (e.g., png, gif, psd, jpeg, and / or jpg) indicating that it contains the image (120).
[0028] According to one embodiment, the electronic device (101) may apply a three-dimensional visual effect to an image (120). The three-dimensional visual effect may be defined to represent dynamic movement based on the image (120). Applying the three-dimensional visual effect to the image (120) may include an operation of generating a video (or information representing the video) in which the positional relationship between an object (152) and a background object of the image (120) gradually changes. The object (152) is a main subject expressed through the image (120), and an area of the image (120) in which the object (152) is displayed may be referred to as a region of interest (ROI), a subject area, and / or a foreground area. The object (152) may be referred to as a salient object and / or a foreground object. The three-dimensional visual effect may be referred to as a parallax effect.
[0029] For example, a user watching a video in which the positional relationship between an object (152) and a background object gradually changes may perceive that the object (152) is positioned closer to the user than the background object. This perception may cause the user to experience a sense of depth and / or three-dimensionality.
[0030] For example, a three-dimensional visual effect applied to an image (120) may be applied to the image (120) to express movement of the background area relative to the object by changing the movement directions and / or movement distances of each of an object (152) of the image (120) (or a foreground area of the image (120) where the object (152) is displayed) and a background object of the image (120) (or another area of the image (120) different from the foreground area and / or the remaining area). The result of applying the three-dimensional visual effect to the image (120) may be referred to as a three-dimensional photograph, a parallax image, and / or a video.
[0031] Referring to FIG. 1, in a state (191) of displaying a screen including an image (120), the electronic device (101) may display a visual object (130) (e.g., a button including designated text such as “Live effect”) for applying a three-dimensional visual effect to the image (120). The electronic device (101) may display a visual object (135) (e.g., a button including designated text such as “Remaster”) for adjusting the color tone, brightness, and / or contrast of the image (120). Through the visual object (130), the electronic device (101) may receive an input for applying a three-dimensional visual effect to the image (120). The input may include a tap gesture on the visual object (130), a mouse click on the visual object (130), and / or a voice command such as an utterance including a word related to the visual object (130) (e.g., “Apply Live effect”). For example, within a state (191) displaying an image (120), the electronic device (101) may receive an input for applying a three-dimensional visual effect to the image (120) according to the actuation of a visual object (130).
[0032] In one embodiment, an electronic device (101) that receives an input for applying a three-dimensional visual effect to an image (120) may identify segmentation information representing an object in the image (120). The segmentation information may include map information (e.g., a segmentation map) indicating whether each pixel included in the image (120) is related to an object. The operation of identifying the segmentation information may include an operation of obtaining the segmentation information by executing a model trained based on artificial intelligence (e.g., by the electronic device (101) and / or a server connected to the electronic device (101). However, the embodiment is not limited thereto, and the operation of identifying the segmentation information may include an operation of directly obtaining or extracting segmentation information stored together with the image (120) (e.g., stored in metadata of the image (120)).
[0033] For example, an electronic device (101) that identifies an object including an edge of an image (120) from segmentation information corresponding to an image (120) may display a first video (150) representing a background area within the image that is moved a first distance beyond the object (152) as a result of applying a stereoscopic visual effect on the display (110). Referring to FIG. 1, an exemplary state (192) of an electronic device (101) displaying the first video (150) as a result of applying a stereoscopic visual effect is illustrated. In order to display a result of applying a stereoscopic visual effect to the image (120), the electronic device (101) may replace or change the image (120) displayed on the display (110) with the first video (150).
[0034] The electronic device (101) can distinguish a visual effect to be applied to the image (120) based on the location of the object (152) (or the area of the image (120) where the object is displayed) within the image (120). For example, if a foreground object is spaced apart from the edge of another image, different from the image (120) of FIG. 1, the electronic device can generate a second video representing the background area that is moved by a shorter distance than the distance by which the background area is moved within the first video (150). For example, since the background area is moved by a second distance that is shorter than the first distance, the second video corresponding to the other image can represent an object that moves relatively slowly (or by a relatively shorter distance) than the background area of the first video.
[0035] Referring to an exemplary state (192) of FIG. 1, the electronic device (101) may display, on the display (110), a result of applying a stereoscopic visual effect (e.g., the first video (150)), together with visual objects (142, 144) for sharing and / or storing the result. For example, the electronic device (101) may display a visual object (142) for sharing and / or transmitting the first video (150), which represents an image (120) to which a stereoscopic visual effect has been applied, on the display (110). The visual object (142) may include, but is not limited to, a designated text such as “share.” The electronic device (101) may display a visual object (144) for storing the image (120) to which the stereoscopic visual effect has been applied, on the display (110). The visual object (144) may include, but is not limited to, designated text such as "save copy." In addition to the visual objects (142, 144), the electronic device (101) may further display a visual object (e.g., a button including designated text such as "save as wallpaper") for setting the video (150) as the background screen of the electronic device (101).
[0036] Referring to FIG. 1, in a state (192) in which a visual object (144) indicating storage of a result of applying a stereoscopic visual effect including a first video (150) is displayed, the electronic device (101) can receive an input for the visual object (144). Based on receiving the input for the visual object (144), the electronic device (101) can store the first video (150) displayed on the display (110) as the result.
[0037] According to one embodiment, the electronic device (101) may select or determine a stereoscopic visual effect to be applied to the image (120) by using the position of an object within the image (120) in which the foreground object (152) is expressed, the size of the object within the image (120), and / or the depth distribution within the image (120). For example, when another object positioned closer to the camera than the foreground object (152) is captured at the time of capturing the image (120), the motion of the foreground object (152) caused by the stereoscopic visual effect may express an unnatural state of obscuring the other object. For example, when both the foreground object (152) and the shadow generated from the foreground object (152) are captured, the motion of the foreground object (152) caused by the stereoscopic visual effect may express the foreground object (152) being separated from the shadow.
[0038] According to one embodiment, the electronic device (101) may change or determine a stereoscopic visual effect to be applied to the image (120) and / or properties of the stereoscopic visual effect using information related to the foreground object (152). For example, since a stereoscopic visual effect suitable for the characteristics of the foreground object (152) is applied to the image (120), the electronic device (101) may generate or display a video (e.g., a first video (150)) that expresses the natural motion of the foreground object (152). For example, when another object is captured closer to the camera than the foreground object (152) at the time of capturing the image (120), the electronic device (101) may apply a stereoscopic visual effect to the image (120) that prevents the foreground object (152) from obscuring the other object, thereby expressing only the natural motion (or motion that follows the laws of physics) of the foreground object (152). By expressing only the natural motion of the foreground object (152), the electronic device (101) can enhance the sense of depth, sense of three-dimensionality, and / or sense of immersion of a user viewing the result of applying a three-dimensional visual effect to an image (120) (e.g., the first video (150)).
[0039] The present disclosure may relate to an electronic device (101) that applies a stereoscopic visual effect to an image (120), and / or changes or determines properties of the stereoscopic visual effect. An operation of the electronic device (101) that applies a stereoscopic visual effect to an image (120) is described with reference to FIG. 3. An operation of the electronic device (101) that selects a stereoscopic visual effect to be applied to an image (120), or changes or determines properties of the stereoscopic visual effect is described with reference to FIG. 4 and / or FIG. 5. Examples of the electronic device (101) identifying each of exemplary images and applying different stereoscopic visual effects are described with reference to FIGS. 6 to 9. An exemplary operation of the electronic device (101) that identifies a foreground object (152), which is a criterion for applying a stereoscopic visual effect, from an image (120) is described with reference to FIG. 10. Stereoscopic visual effects applicable to an image (120) by an electronic device (101) are described with reference to FIG. 11A and / or FIG. 11B. An operation by which the electronic device (101) receives a user input for changing the properties of a stereoscopic visual effect is described with reference to FIG. 12. An exemplary operation by which the electronic device (101) generates depth information for an image (120), which is used to apply a stereoscopic visual effect to the image (120), is described with reference to FIG. 13.
[0040] FIG. 2 is a block diagram of an electronic device (101) according to one embodiment. Referring to FIG. 2, the electronic device (101) may be one of various forms of electronic devices, such as a laptop PC (personal computer) (290), smartphones (291) having various form factors (e.g., a bar-type smartphone (291-1), a foldable-type smartphone (291-2) having the appearance of the electronic device (101) of FIG. 1, or a sliderable (or rollable) type smartphone (291-3)), a tablet PC (292), a head-mounted display (HMD) device (293), a watch (294), a cellular phone (not shown), and other similar computing devices (not shown).
[0041] In one embodiment, the electronic device (101) may be referred to as a mobile device, a user equipment (UE) (or user terminal), a multi-function device, a portable communication device, a portable device, or a server. The form factor of the electronic device (101) is not limited to the exemplary form factors illustrated in FIG. 2. For example, in some embodiments, the electronic device (101) may be included as an electronic control unit (ECU) in a vehicle (e.g., an electric vehicle (EV)). For example, the electronic device (101) may have a form suitable for displaying images and / or videos.
[0042] Referring to FIG. 2, according to one embodiment, an electronic device (101) may include a processor (210) and a memory (220). The electronic device (101) may further include a display (110). The processor (210) may be electrically and / or operatively coupled with the memory (220) and / or the display (110). Electrical coupling of the electronic components may include a state in which a wired signal path (or a connection for wireless communication) for transmitting signals is established between the electronic components. Operationally coupling of the electronic components may include a state in which the electronic components are directly coupled (or a state in which the electronic components are indirectly coupled) such that one of the electronic components controls another electronic component. Referring to FIG. 2, an electrical connection between the processor (210), the memory (220), and the display (110) based on electronic components, referred to as a communication bus (202), is schematically illustrated. Through a communication bus (202), the processor (210), memory (220), and display (110) can be communicatively coupled.
[0043] Referring to FIG. 2, a processor (210) of an electronic device (101) may include circuits (e.g., processing circuits and / or cores) for performing operations (e.g., arithmetic operations and / or logical operations) on data. Binary codes (e.g., instructions) representing the operations may be input to the processor (210). The processor (210) may include a central processing unit (CPU), a graphic processing unit (GPU), and / or a neural processing unit (NPU). The processor (210) may be referred to as an application processor (AP) and / or a system on a chip (SoC). The processor (210) may have a structure (e.g., a multi-core structure based on a combination of multiple core circuits such as a dual core, a quad core, a hexa core, or an octa core) for loading (or fetching) and / or executing multiple instructions simultaneously. Within an electronic device (101) comprising at least one processor, including a processor (210), the at least one processor may individually or collectively perform the operations of the present disclosure. For example, the at least one processor may individually and / or collectively perform the operations of FIGS. 3 to 5 and / or FIG. 10 by executing instructions stored in the memory (220).
[0044] The memory (220) of FIG. 2 may include a circuit for storing data (or instructions) input to or output from the processor (210). The memory (220) may include volatile memory, such as random-access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM). The non-volatile memory may be referred to as storage. The volatile memory may include, for example, at least one of dynamic RAM (DRAM), static RAM (SRAM), cache RAM, and pseudo SRAM (PSRAM). The non-volatile memory may include, for example, at least one of programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), flash memory, hard disk, compact disc, solid state drive (SSD), and embedded multimedia card (eMMC). The memory (220) may include one or more storage media (e.g., the volatile memory and / or non-volatile memory described above) located in a distributed manner in the electronic device (101). The processor (210) of the electronic device (101) may execute instructions of the memory (220) within the electronic device (101) to perform functions and / or operations indicated by the instructions (e.g., the operations of FIGS. 3 to 5 and / or 10).
[0045] The display (110) of the electronic device (101) may include a circuit for visualizing information provided from the processor (210). The display (110) may include a liquid crystal display (LCD), a plasma display panel (PDP), and / or light emitting diodes (LEDs). The LEDs may include organic LEDs (OLEDs). However, the embodiment is not limited thereto, and the display (110) may include electronic paper. The display area (or active area) of the display (110) may include an area where light is emitted, formed by pixels (e.g., activated pixels) of the display (110). The display (110) may include a sensor (e.g., a touch sensor) for detecting an external object (e.g., a user's finger) on the display (110). The sensor may be included in the display (110) in the form of a panel (e.g., a touch sensor panel (TSP)).
[0046] Referring to FIG. 2, programs (e.g., a salient object detector (241) (or a notable object detector), a depth detector (242), a depth inversion detector (243), a float detector (244), a pixel unprojector (245), and / or a 3D renderer (246)) executed by a processor (210) to process an image (230) (e.g., image (120) of FIG. 1) are illustrated. The programs may be independently installed in memory (220) (e.g., a collection of instructions and resources, referred to as a package), or may be stored in memory (220) as sub-routines of a single program (or an applet or a dynamic link library (DLL)).
[0047] According to one embodiment, the processor (210) of the electronic device (101) may execute a primary object detector (241) to detect a foreground object (e.g., a subject) associated with an image (230). By executing the primary object detector (241), the processor (210) may segment a region associated with the object within the image (230). For example, the processor (210) may obtain or generate information indicating the location, size, and / or shape of the region associated with the object. To obtain the information, the primary object detector (241) may include an artificial intelligence model trained to output the information from a color distribution of the image (230). The artificial intelligence model is a computational model designed to simulate neural activity (or cognitive activity) of a living organism, and may include a program, hardware (e.g., an NPU and / or a GPU), or any combination thereof for performing computations represented by the computational model. For example, in one embodiment, an artificial intelligence model may be trained on color distributions of images having objects having known locations, sizes, and / or shapes from the images. The primary object detector (241) may input the color distribution of the image (230) into the trained artificial intelligence model, and may receive information about the location, size, and / or shape of a region associated with an object within the image (230) as output from the trained artificial intelligence model.
[0048] According to one embodiment, the processor (210) of the electronic device (101) may execute a depth information estimator (242) to calculate or estimate depth values corresponding to each of the pixels of an image (230) representing a two-dimensional color distribution. The set of depth values and / or the two-dimensional distribution of the depth values may be referred to as depth information and / or a depth map. The depth value corresponding to each pixel may represent a distance (or a relative value of the distance) between a subject corresponding to the pixel and a camera used to capture the image (230). The depth information estimator (242) may include a program for decoding and / or extracting depth information included in metadata of the image (230) (e.g., sensor data of a time-of-flight (ToF) camera and / or a light detection and ranging (LiDAR) camera that was operating at the time of capturing the image (230). For example, depth information may be represented by sensor data of a ToF sensor and / or a LiDAR sensor acquired together with the image (230). The depth information estimator (242) may output depth values corresponding to each pixel of the image (230), or may include a computational model (e.g., an artificial intelligence model) trained to estimate depth information. For example, in one embodiment, the artificial intelligence model may be trained by pixels of images for which the distance (or the relative value of the distance) between a subject corresponding to a pixel and a camera used to capture the image is known. The depth information estimator (242) may input pixel values of the image (230) into the trained artificial intelligence model, and may receive estimated depth information from the artificial intelligence model as an output.
[0049] According to one embodiment, the processor (210) of the electronic device (101) may execute a depth inversion detector (243) to detect or identify, from depth information corresponding to the image (230), another object (e.g., a subject included in a background area) located closer to the camera that captured the image (230) than the foreground object (e.g., the foreground object (152) of FIG. 1). The depth inversion detector (243) may compare depth values of an object in the image (230) detected by the main object detector (241) with depth values of a remaining area of the image (230) that is different from the area in which the object is displayed (e.g., a background area), to detect or identify an external object located between the foreground object and the camera at the time the image (230) was captured.
[0050] According to one embodiment, the processor (210) of the electronic device (101) may execute a floating object detector (244) to determine the location and / or size of an object within an image (230). For example, the processor (210) may calculate or identify a distance between the object and an edge of the image (230) (e.g., the bottom of the image (230)). For example, the processor (210) may identify whether the object is spaced from the edge of the image (230). The processor (210), having determined the distance between the edge of the image (230) and the object, may determine or determine whether the distance exceeds a specified threshold. The distance between the edge of the image (230) and the object (e.g., the minimum value of the distance between the pixels included in the area where the object is displayed and the edge) determined by executing the floating object detector (244) may be used to determine (or select) a stereoscopic visual effect to be applied to the image (230). The processor (210) can execute a floating object detector (244) to determine the distance an object (e.g., a salient object) is spaced from the bottom of the image (230). The processor (210) can determine whether the object is spaced from the bottom by more than a reference distance (e.g., a reference distance for changing a stereoscopic visual effect).
[0051] According to one embodiment, the processor (210) of the electronic device (101) may obtain or determine three-dimensional coordinates (e.g., spatial coordinates) corresponding to pixels included in the image (230) by executing the pixel unprojector (245). The three-dimensional coordinates may be referred to as vertices (or voxels) representing three-dimensional points within a virtual space. By executing the pixel unprojector (245), the processor (210) may obtain or generate a three-dimensional distribution (e.g., point cloud) of pixels included in the image (230). The processor (210) executing the pixel unprojector (245) may generate or determine three-dimensional coordinates of each pixel of the image (230) by using depth information (e.g., depth map) generated by the depth information estimator (242).
[0052] According to one embodiment, the 3D renderer (246) of the electronic device (101) may perform stereoscopic rendering of an image (230) based on 3D coordinates generated by the pixel unprojector (245). For example, the processor (210) may place a virtual camera within a virtual space including the 3D coordinates. Stereoscopic rendering of the image (230) may include an operation of generating an image representing a view of the virtual space as seen from the virtual camera. By executing the 3D renderer (246), the processor (210) may render or generate a 2D image from the 3D scene generated by the pixel unprojector (245). For example, the processor (210) may render the 2D image using the viewpoint and / or position of the virtual camera.
[0053] According to one embodiment, the processor (210) may apply a stereoscopic visual effect to the image (230). For example, based on the stereoscopic visual effect, the 3D coordinates generated by executing the pixel unprojector (245) may be at least partially changed. For example, the processor (210) may move the 3D coordinates of pixels corresponding to an object differently from the 3D coordinates of pixels corresponding to a background area. For example, the processor (210) may move the coordinates of a virtual camera formed within the virtual space to a direction and / or position related to the stereoscopic visual effect. While moving the 3D coordinates of the pixels and / or the 3D coordinates of the virtual camera, the processor (210) may perform rendering based on the 3D renderer (246) to generate or obtain an animation and / or video (e.g., the first video (150) of FIG. 1) expressing the image (230) to which the stereoscopic visual effect is applied.
[0054] In one embodiment, the electronic device (101) may analyze the image (230) to vary or determine the degree to which a stereoscopic visual effect (or parallax effect) is applied. For example, the electronic device (101) may vary or determine the degree to represent natural motion between a salient object and a background object in the image (230).
[0055] The present disclosure describes, but is not limited to, an operation of generating an animation and / or video representing an image (230) to which a stereoscopic visual effect has been applied. For example, in one embodiment where the electronic device (101) is an HMD device (293), the HMD device (293) may generate or display a spatial video and / or spatial image representing a stereoscopic motion of a background area for an object by applying a stereoscopic visual effect to the image (230). For example, a point cloud rendered by a pixel unprojector (245) and corresponding to pixels of the image (230) may be displayed stereoscopically to a user wearing the HMD device (293) based on binocular parallax. For example, the HMD device (293) can project images and / or videos having the binocular parallax to each of the two eyes of a user wearing the HMD device (293) by using the binocular parallax corresponding to the depth values of the pixels of the image (230).
[0056] As described above, according to one embodiment, the electronic device (101) can select or determine a stereoscopic visual effect to be applied to the image (230) by using not only the depth information corresponding to the image (230) measured by the depth information estimator (242), but also the information identified by the primary object detector (241), the depth inversion detector (243), and / or the floating object detector (244). By using the depth information and additional information, the electronic device (101) can segment or extract objects from the image (230) more accurately. The electronic device (101) can select a stereoscopic visual effect suitable for the image (230) and apply the selected stereoscopic visual effect to the image (230). By applying the stereoscopic visual effect suitable for the image (230), the electronic device (101) can provide an immersive user experience for the result of applying the stereoscopic visual effect.
[0057] Hereinafter, exemplary operations of an electronic device (101) and / or a processor (210) for applying a three-dimensional visual effect to an image (230) are described with reference to FIG. 3.
[0058] FIG. 3 illustrates a flowchart of an electronic device according to one embodiment. The electronic device of FIG. 3 may include the electronic device (101) of FIG. 1 and / or FIG. 2. The operations of FIG. 3 may be performed by the electronic device (101) and / or the processor (210) of FIG. 2. The operations of FIG. 3 may be performed based on the execution of programs illustrated in FIG. 2 (e.g., a salient object detector (241), a depth information estimator (242), a depth inversion detector (243), a floating object detector (244), a pixel unprojector (245), and / or a 3D renderer (246)).
[0059] The order in which the operations of FIG. 3 are performed may vary depending on the embodiment. For example, according to one embodiment, the electronic device may perform the operations of FIG. 3 in a different order than that shown in FIG. 3, or may perform at least two operations substantially simultaneously.
[0060] Referring to FIG. 3, within operation (310), according to one embodiment, a processor of an electronic device may receive an input for applying a stereoscopic visual effect to an image (315). The input may include an input indicating selection of a visual object (130), as described with reference to FIG. 1. The input may be detected based on a touch gesture on a display (e.g., display (110) of FIG. 1 and / or FIG. 2), a click received via a mouse, a user's gaze identified via a sensor (e.g., an eye-tracking camera (ET-CAM)) (e.g., a user's gaze wearing an HMD device (293) of FIG. 2), and / or a user's speech.
[0061] Referring to FIG. 3, in operation (320), according to an embodiment, a processor of an electronic device may obtain depth information (325) corresponding to an image (315). Operation (320) may be performed by executing the depth information estimator (242) of FIG. 2. The processor may obtain depth information (325) including depth values each corresponding to pixels of the image (315). According to an embodiment, the processor may obtain depth information (325) from metadata corresponding to the image (315). The depth information (325) may be included in a file including the image (315), for example, according to the exchangeable image file format (EXIF). According to an embodiment, when applying a stereoscopic visual effect to the image (315), the processor of the electronic device may further utilize characteristics of content included in the image (315) in addition to the depth information (325) obtained based on operation (320) in order to provide a natural stereoscopic visual effect.
[0062] Referring to FIG. 3, in operation (330), according to one embodiment, a processor of an electronic device may identify an object (332) in an image (315). The processor may perform operation (330) by executing a salient object detector (241). The processor may obtain or generate information indicating a probability that each pixel of the image (315) corresponds to an object. The information may be referred to as segmentation information, a segmentation map, salient information, and / or salient map. The probability may be obtained for each pixel of the image (315) using an artificial intelligence model for object recognition. For example, the processor may obtain segmentation information corresponding to the image (315) using the aforementioned model for object recognition using artificial intelligence. Using the segmentation information, the processor may identify a location, size, and / or shape of the object (332) in the image (315).
[0063] Referring to FIG. 3, in operation (340), according to one embodiment, the processor of the electronic device may obtain a background area (344) by combining the inpainting area (334) that replaces the object (332) and the remaining area of the image excluding the object. For example, the background area (344) may include the inpainting area (334) and the remaining area of the image (315) that is different from the object. For example, the processor may perform inpainting on the object (332) to obtain the inpainting area (334) that replaces the object. The inpainting area (334) may correspond to a portion where the object (332) is located.
[0064] In one embodiment, the inpainting for the object (332) may be obtained from a computational model (e.g., a computational model referred to as generative artificial intelligence) that is fed with an image (315) (or another image in which the inpainting region (334) is filled with a specified color, such as black). The computational model may be an artificial intelligence model that is trained to determine colors of pixels in the inpainting region (334) based on the content of the remainder of the image (315) that is different from the inpainting region (334). The computational model may further receive a natural language sentence (e.g., a prompt) describing the content to be expressed through the inpainting region (334). The artificial intelligence model may be executed by the electronic device (e.g., an on-device model) or may be executed by an external electronic device (e.g., a server) connected to the electronic device via communication circuitry. The above artificial intelligence model can be trained to generate an inpainting region (334) that matches the background region (344) of the image (315) by receiving a prompt generated based on the colors of pixels of the image (315) adjacent to the object (332) and / or the analysis results of the image (315).
[0065] Referring to FIG. 3, in operation (350), a processor of an electronic device according to an embodiment may perform three-dimensional rendering of an object (332) and a background area (344). The processor may execute a pixel unprojector (245) and / or a 3D renderer (246) of FIG. 2 to perform operation (350). Referring to FIG. 3, the processor may generate a virtual space including the object (332) and the background area (344). The pixels of the image (315) may be arranged in the virtual space according to depth values corresponding to each pixel. Referring to FIG. 3, the depth values of the pixels corresponding to the object (332) may be smaller than or may exhibit a closer depth than the depth values of the pixels corresponding to the background area (344).
[0066] Referring to FIG. 3, the processor may perform rendering of the operation (350) using a virtual camera (356) positioned within the virtual space. However, the embodiment is not limited to rendering based on the virtual camera (356), and the rendering of the operation (350) may be performed based on the view of a virtual user. For example, the rendering of the operation (350) may include generating an image and / or video representing a view of the virtual space as seen from the virtual camera (356). Applying a stereoscopic visual effect may include generating a video and / or animation representing the motion of the background area (344) relative to the object (332) by gradually changing the position of at least one of the object (332), the background area (344), or the virtual camera (356) within the virtual space. In one embodiment, the electronic device may (automatically) generate or suggest a stereoscopic visual effect and / or movement of the virtual camera (356) to be applied to the image (315). For example, the electronic device may display a user interface (UI) for changing the direction, speed, and / or trajectory of movement of the virtual camera (356). Through the UI, the electronic device may receive input from the user for changing the settings of the virtual camera (356) and / or the stereoscopic visual effect. Based on the input, the electronic device may perform rendering of the motion (360).
[0067] Referring to FIG. 3, in operation (360), according to one embodiment, a processor of an electronic device may display a video (e.g., the first video (150) of FIG. 1) representing the result of performing three-dimensional rendering. The video may represent an image (315) to which a stereoscopic visual effect is applied. To apply the stereoscopic visual effect, information representing different stereoscopic visual effects may be stored in a memory of the electronic device (e.g., the memory (220) of FIG. 2). The processor may identify the information representing the stereoscopic visual effects from the memory. The information representing the stereoscopic visual effect may be defined for key frames of the video of operation (360), which is the result of applying the stereoscopic visual effect. A key frame may be described as a reference frame for other adjacent frames in the time domain within a sequence of images (or image frames) included in the video. For example, in the time domain, the colors of pixels of other frames adjacent to a key frame may be set to a difference value with respect to the colors of pixels of the key frame.
[0068] For example, information representing a stereoscopic visual effect may indicate the position of at least one of an object (332), a background area (344), and a virtual camera (356) in each of a plurality of key frames included in the video. For example, the information may include a horizontal position of a first layer corresponding to the background area (344) (e.g., a position on the y-axis in FIG. 3), a horizontal position of a second layer corresponding to the object (332) (e.g., a position on the y-axis in FIG. 3), and a horizontal position of a virtual camera (356) moved for three-dimensional rendering within a virtual space including the first layer and the second layer (e.g., a position on the y-axis in FIG. 3). The horizontal position may be expressed as a coordinate value of the y-axis. The information may be expressed in a format such as JSON (JavaScript object notation). Exemplary information loaded to apply a stereoscopic visual effect is described with reference to FIG. 6.
[0069] Below, with reference to FIG. 4, the operation of an electronic device that determines the motion of an object (332), a background area (344), and a virtual camera (356) when performing rendering of operations (350, 360) is described.
[0070] FIG. 4 illustrates an operation of an electronic device that determines a stereoscopic visual effect to be applied to an image by using a location of an object and / or a depth distribution of an image including the object, according to one embodiment. The electronic device of FIG. 4 may include the electronic device (101) of FIG. 1 and / or FIG. 2. The operations of FIG. 4 may be performed based on the execution of programs (e.g., a salient object detector (241), a depth information estimator (242), a depth inversion detector (243), a floating object detector (244), a pixel unprojector (245), and / or a 3D renderer (246)) illustrated in FIG. 2. The operations of FIG. 4 may be performed by the electronic device (101) and / or the processor (210) of FIG. 2.
[0071] The order in which the operations of FIG. 4 are performed may vary depending on the embodiment. For example, according to one embodiment, the electronic device may perform the operations of FIG. 4 in a different order than that shown in FIG. 4, or may perform at least two operations substantially simultaneously.
[0072] Referring to FIG. 4, in operation (410), a processor of an electronic device according to one embodiment may detect or determine whether an object has moved away from the edge of an image (e.g., image (120) of FIG. 1, image (230) of FIG. 2, and / or image (315) of FIG. 3). In one embodiment, the electronic device may determine whether the object has moved away from the bottom of the image (315) based on whether the object has moved away from the edge of the image (315) by a threshold distance, in order to detect whether the object has moved away from the edge of the image (315). The threshold distance may be preset. In operation (410), the processor may determine whether the object is floating. Floating of the object may be described as a state in which a foreground object expressed through the object is not cropped by the field of view (FoV) (or angle of view) of a camera that captured the image. For example, an object may be described as floating if the entire foreground object is represented through the image, or if the object is separated from an edge, such as the bottom of the image.
[0073] The operation (410) of FIG. 4 may be performed by executing the main object detector (241) of FIG. 2. The processor that detects the object may perform the operation (410) using the location, shape, and / or size of the object within the image. For example, the processor may identify the minimum value of the y-coordinate values of the pixels included in the object. If the minimum value is different from the y-coordinate value corresponding to the bottom of the image, the processor may determine that the object is spaced away from the edge of the image. However, the embodiment is not limited thereto, and the processor may determine that the object is spaced away from the edge of the image if the object is spaced away from the edge of the image by more than a specified threshold.
[0074] For example, the processor may perform operation (410) by executing the pseudo code of Table 1.
[0075]
[0076] Referring to Table 1, the processor is a list (e.g., List <object>For each of the foreground objects included in the salientObjects, the operation (410) may be repeatedly performed. The foreground object may include a face, a person, and / or an animal that may be identified as a main object. Referring to Table 1, it may be determined whether an object corresponding to a foreground object is spaced from the edge of the image (e.g., the bottom of the image) based on a threshold (e.g., THRESHOLD_FLOATING_OBJECT_TO_BOTTOM_RATIO). The threshold may be preset. Referring to Table 1, when multiple foreground objects are identified from the image, the determination may be made by comparing the threshold with the distance between the object located at the bottommost position in the image, among the objects corresponding to each of the multiple foreground objects, from the edge and the ratio between the size of the image (maxBottomDistance / segmentationMap.getHeight()). The distance from the edge of the object may be obtained by executing the function (calculateBottomDistance) of the pseudo code of Table 1. For example, if the threshold is 0.09, then if the ratio exceeds 9%, the object can be determined to be away from the edge of the image.
[0077] Referring to FIG. 4, if the object is spaced from the edge of the image (operation (410) - Yes), the processor may perform operation (420). If the object is not spaced from the edge of the image (e.g., the object is in contact with the edge, or the object includes the edge (operation (410) - No), the processor may perform operation (430).
[0078] Referring to FIG. 4, within operations (420) and / or (430), a processor of an electronic device according to an embodiment may determine or determine whether a depth inversion of a background region with respect to an object has been detected. For example, a depth inversion may indicate a case in which at least one portion of an additional object and / or background region exists that has a shallower depth than the depth of the object. Operations (420, 430) may be performed by executing the depth information estimator (242) and / or the depth inversion detector (243) of FIG. 2. An example of an operation for detecting depth inversion in operations (420, 430) is described with reference to FIG. 5.
[0079] To detect depth inversion of actions (420, 430), the processor may identify depth information corresponding to the image of action (410). If, from the depth information, the processor identifies at least a portion of the background area that appears to have a shallower distance than the distance between the object and the camera that captured the image, the processor may determine that depth inversion has been detected. If, from the depth information, the entire background area has a greater distance than the distance between the object and the camera, the processor may determine that depth inversion has not been detected.
[0080] Referring to FIG. 4, within operation (420), if depth inversion is detected (420 - Yes), the processor may perform operation (425). Within operation (420), if depth inversion is not detected (420 - No), the processor may perform operation (440). Within operation (430), if depth inversion is detected (430 - Yes), the processor may perform operation (425). Within operation (430), if depth inversion is not detected (430 - No), the processor may perform operation (450).
[0081] Referring to FIG. 4, in operation (450), according to one embodiment, a processor of an electronic device may apply a visual effect for dynamic action among stereoscopic visual effects as an image. The processor may identify information representing the stereoscopic visual effect. Without changing or reducing the information, the processor may apply the stereoscopic visual effect represented by the information as an image. For example, the information may be represented in the JSON format of Table 2.
[0082]
[0083] Referring to Table 2, a JSON object with the name "keyframes" can be defined. The square brackets of lines 1 and 4 can indicate that the JSON object with the name "keyframes" contains numeric values between the square brackets. Line number 2 can indicate the horizontal position of the background area ("layer_position[0].x": -0.09), the horizontal position of the object ("layer_position[1].x": 0.01), the horizontal position of the virtual camera ("camera_eye_x": -0.1), and the direction of the virtual camera ("camera_dir_x": 0.1) in the key frame corresponding to time = 0.0. The key frame corresponding to time = 0.0 can correspond to the time when the animation starts. For example, time = 0.0 can correspond to the time when the animation starts. For example, time = 0.0 may represent the time when an animation starts in which a virtual camera moves relative to an object to indicate motion of a background area. Line number 3 may represent the horizontal position of the background area ("layer_position[0].x": 0.09), the horizontal position of the object ("layer_position[1].x": -0.01), the horizontal position of the virtual camera ("camera_eye_x": 0.1), and the direction of the virtual camera ("camera_dir_x": -0.1) at a key frame corresponding to time = 1.0 (e.g., another key frame after the key frame corresponding to time = 0.0). However, the embodiment is not limited thereto, and the direction of the background area ("layer_direction[0].x") may be defined to change according to the key frame. The key frame corresponding to time = 1.0 may correspond to a key frame to be played at the time when the animation ends. For example, time = 1.0 may correspond to the time when the animation ends.The orientation of the virtual camera may represent the azimuth and / or rotation angle of the virtual camera.
[0084] According to one embodiment, not only the movement of the background area along the x-axis (e.g., the movement of the "layer_position[0].x" value) but also the direction of the virtual camera (e.g., the movement of the "camera_dir_x" value) are defined, so that the electronic device can generate an animation in which the motion of the virtual camera and / or the motion of the background area is expressed dynamically with respect to the object. Although an embodiment has been described in which the background area, object, and / or virtual camera are moved based on a horizontal position, the embodiment is not limited thereto, and the information in Table 2 may be defined such that the background area, object, and / or virtual camera are moved along the vertical axis.
[0085] The "time" in Table 2 is a relative value indicating the point in time of the corresponding key frame, and for example, when generating a 4-second video, it can indicate the temporal position of the key frame within the time interval from 0 seconds (time = 0.0) to 4 seconds (time = 1.0). For example, the key frame corresponding to time = 0.0 can correspond to an image to be displayed at 0 seconds of the video that expresses an image with a stereoscopic visual effect applied. For example, the key frame corresponding to time = 1.0 can correspond to an image to be displayed at 4 seconds of the video. The processor can use the information defined as in Table 2 to gradually move objects, background areas, and virtual cameras within a virtual space, thereby rendering image frames included in the 4-second video.
[0086] Referring to FIG. 4, in operation (425), according to one embodiment, a processor of an electronic device may apply a static visual effect among stereoscopic visual effects to an image. For example, the processor may reduce a deviation in horizontal positions of a background region indicated by information defined as in Table 2. For example, the processor may at least partially change a variable of the information (e.g., layer_position[0].x) such that the deviation in horizontal positions of the background region across key frames is reduced. As the horizontal position of at least one of the key frames changes, the degree to which the background region moves may be reduced. The processor may generate a video including a background region that moves along a horizontal position with a reduced deviation, thereby representing an image to which a static stereoscopic visual effect is applied.
[0087] Referring to FIG. 4, in operation (440), a processor of an electronic device according to an embodiment may apply a visual effect for a jump action among stereoscopic visual effects as an image. Based on the dynamic characteristics of the jump action, the processor may apply a stereoscopic visual effect that emphasizes the dynamic characteristics among the stereoscopic visual effects as an image. When applying a stereoscopic visual effect based on operation (440) as an image, the processor may not adjust the deviation of the horizontal positions of the background area, unlike operation (425).
[0088] Referring to Figure 4, the stereoscopic visual effects applied to the image, depending on whether the object is floating and / or whether depth inversion is detected, can be summarized as in Table 3.
[0089]
[0090] For example, a processor that identifies at least a portion of a background region, from the depth information, having a shallower depth than the depth between the object and the camera that captured the image, may generate or display a video representing an image having a limited stereoscopic visual effect applied thereto. For example, based on identifying an object, from the depth information, having a shallower depth than the depth of the background region, information indicating the position and / or size of the object (e.g., segmentation information) may be further utilized to select or determine a stereoscopic visual effect to be applied to the image.
[0091] FIG. 5 illustrates the operation of an electronic device for comparing depths of an object and a background region, according to one embodiment. The electronic device of FIG. 5 may include the electronic device (101) of FIG. 1 and / or FIG. 2. The operations of FIG. 5 may be performed based on the execution of the program (e.g., the depth inversion detector (243)) illustrated in FIG. 2. The operations of FIG. 5 may be performed by the electronic device (101) and / or the processor (210) of FIG. 2.
[0092] The order in which the operations of FIG. 5 are performed may vary depending on the embodiment. For example, according to one embodiment, the electronic device may perform the operations of FIG. 5 in a different order than that shown in FIG. 5, or may perform at least two operations substantially simultaneously.
[0093] Referring to FIG. 5, in operation (510), a processor of an electronic device according to one embodiment may identify depth information and segmentation information corresponding to an image (e.g., image (120) of FIG. 1, image (230) of FIG. 2, and / or image (315) of FIG. 3). The depth information may include a two-dimensional array of depth values corresponding to each pixel of the image. The segmentation information may include a two-dimensional array of probabilities that each pixel of the image corresponds to a foreground object. The segmentation information may include a saliency map. The saliency map may include a two-dimensional image that highlights areas within the image that are likely to be preferentially focused on by a user and / or areas that are likely to be recognized by a machine learning model. For example, pixels in the saliency map may indicate the importance of corresponding pixels in the image. Within operation (510), the processor may perform resizing of the depth information and segmentation information so that the depth information and segmentation information have the same size. For example, the resizing may include normalization.
[0094] Referring to FIG. 5, in operation 520, a processor of an electronic device according to an embodiment may obtain an average depth value of an object in an image, indicated by segmentation information. The object may include pixels having a probability exceeding a specified threshold probability. The electronic device may obtain or determine a bounding box having a shape surrounding the object. The bounding box may be expanded in the x-axis and y-axis directions by a margin (e.g., X_MARGIN and / or Y_MARGIN in Table 4 below). The processor may calculate an average depth value for pixels of the object included in the expanded bounding box. The processor may apply a Gaussian blur (or Gaussian filter) to the depth values of the pixels to reduce errors included in the depth values. By calculating an average of the depth values having the reduced error, the processor may obtain or calculate the average depth value of operation 520.
[0095] Referring to FIG. 5, in operation (530), the processor of an electronic device according to one embodiment may obtain an average depth value of a background area. The background area may correspond to the remaining area of the image, which is different from the object of operation (520). The average of the depth values of the pixels included in the background area may be determined as the average depth value of operation (530).
[0096] Referring to FIG. 5, in operation (540), a processor of an electronic device according to an embodiment may determine whether an average depth value of a background region exceeds an average depth value of an object. If the average depth value of the background region is less than or equal to the average depth value of the object (540-No), the processor may perform operation (580). If the average depth value of the background region is less than or equal to the average depth value of the object, it may mean that the depth values of the background region are smaller than the depth values of the object on average. If the average depth value of the background region exceeds the average depth value of the object (540-Yes), the processor may perform operation (550).
[0097] Referring to FIG. 5, in operation (550), according to one embodiment, a processor of an electronic device may obtain a maximum depth value of a background area. For example, the processor may obtain or retrieve a maximum value among depth values of pixels included in the background area.
[0098] Referring to FIG. 5, in operation (560), the processor of the electronic device according to one embodiment may determine or verify whether the maximum depth value of the background region is greater than or equal to the average depth value of the object. For example, the processor may compare the maximum depth value of the background region with a combination of the average depth value of the object and a specified tolerance. If the maximum depth value of the background region is greater than or equal to the average depth value of the object, or if the maximum depth value of the background region is equal to or greater than the sum of the average depth value of the foreground object and the tolerance (560—Yes), the processor may perform operation (570). If the maximum depth value of the background region is less than the average depth value of the object, or if the maximum depth value of the background region is less than the combination (560—No), the processor may perform operation (580).
[0099] Referring to FIG. 5, in operation (570), the processor of the electronic device according to one embodiment may determine that depth inversion of a background region for an object has not occurred. For example, the processor may determine that any subject corresponding to the background region is positioned further from the camera than a foreground object corresponding to the object. Based on the determination of operation (570), the processor may perform any one of operations (440, 450) of FIG. 4.
[0100] Referring to FIG. 5, in operation (580), according to one embodiment, the processor of the electronic device may determine that depth inversion has occurred in the background region for an object. For example, the processor may determine that a subject in the background region is positioned closer to the camera than a foreground object corresponding to the object. Based on the determination in operation (580), the processor may perform operation (425) of FIG. 4.
[0101] In one embodiment, the operations of FIG. 5 can be implemented based on the pseudo code of Table 4.
[0102]
[0103] Hereinafter, with reference to FIGS. 6 to 9, exemplary operations for selecting a stereoscopic visual effect to be applied to an image based on an object and / or background area of the image by the processor are described.
[0104] FIG. 6 illustrates the operation of an electronic device applying a stereoscopic visual effect to an exemplary image (610), according to one embodiment. The electronic device of FIG. 6 may include the electronic device (101) of FIG. 1 and / or FIG. 2 . The operation performed by the electronic device of FIG. 6 may be related to at least one of the operations of FIGS. 3 to 5 .
[0105] Referring to Fig. 6, an exemplary image (610) is illustrated. An electronic device can obtain depth information (620) of the image (610). Referring to Fig. 6, the density of dots of the depth information (620) can indicate a distance represented by a depth value of the depth information (620). For example, a higher density of dots of a part of the depth information (620) than a higher density of dots of another part can indicate that the depth values of the part are lower than the depth values of the other part. In other words, a higher density of dots of a part of the depth information (620) can indicate that an object associated with the part is positioned closer to the camera than an object associated with the other part.
[0106] In one embodiment, the electronic device can identify an object (630) in an image (610). The object (630) can be identified using a model trained to process at least one of a depth distribution represented by depth information (620) and / or a color distribution of pixels in the image (610).
[0107] Referring to FIG. 6, an electronic device that has identified an object (630) can infer content beyond the object (630) to generate or obtain a background area (640). The background area (640) can be a combination of the remaining area of the image (610) that is different from the object (630), and content obscured by the object (630) (e.g., an inpainting area).
[0108] As described above with reference to FIGS. 3 to 5, the electronic device can determine a stereoscopic visual effect to be applied to the image (610) by using the location of the object (630) within the image (610). Referring to the location of the object (630) within the image (610) of FIG. 6, at least a portion of the object (630) may contact an edge (e.g., a bottom) of the image (610). A foreground object (e.g., a person) represented by the object (630) may not be completely included in the FoV of the camera that captured the image (610).
[0109] Referring to the depth information (620) of FIG. 6, the depth values of the object (630) may be less than the depth values of the background area (or the remaining area of the image). That is, an electronic device that applies a stereoscopic visual effect with the image (610) and / or the depth information (620) may determine that depth inversion of the background area with respect to the object (630) has not occurred. Since the distance between the object (630) and the camera (or image sensor) is less than the distance between the background area and the camera, the depth corresponding to the object (630) may be shallower than the depth corresponding to the background area. For example, the depth value of the object (630) may be less than the depth value of the background area. If the object (630) includes at least a portion of the edge of the image (610) and depth inversion has not occurred, the processor may apply a stereoscopic visual effect for dynamic action to the image (610) based on the operation (450) of FIG. 4.
[0110] Referring to FIG. 6, image frames (651, 652, 653) of a video are illustrated, which are generated by applying a stereoscopic visual effect to an image (610). The processor may generate, display, or store the video, which is configured to sequentially display the image frames (651, 652, 653) in the time domain. The processor may obtain or generate the image frames (651, 652, 653) by moving and / or rotating an object (630), a background area (640), and a virtual camera (e.g., the virtual camera (356) of FIG. 3) within a virtual space using the information of Table 2. Referring to FIG. 6, image frames (651, 652, 653) sequentially displayed within a time interval between t1 and t3 can be generated so that the background area (640) gradually moves in a specified direction (e.g., toward the left side of the sheet on which the drawing is printed) centered on the object (630).
[0111] FIG. 7 illustrates the operation of an electronic device applying a stereoscopic visual effect to an exemplary image (710), according to one embodiment. The electronic device of FIG. 7 may include the electronic device (101) of FIG. 1 and / or FIG. 2 . The operation performed by the electronic device of FIG. 7 may be related to at least one of the operations of FIGS. 3 to 5 .
[0112] Referring to Fig. 7, an exemplary image (710) is illustrated. An electronic device can identify depth information (720) (or depth map) of the image (710). Referring to Fig. 7, the depth information (720) can represent a two-dimensional distribution of depth values corresponding to pixels of the image (710). The two-dimensional distribution of depth values represented by the depth information (720) is illustrated based on the density of dots in Fig. 7. For example, a portion where dots with a relatively high density are illustrated in the depth information (720) can have a lower depth value than a portion where dots with a relatively low density are illustrated.
[0113] Referring to FIG. 7, an electronic device may obtain or identify information (e.g., segmentation information) representing an object (730) in an image (710). Having identified the object (730), the electronic device may generate or obtain a background region (740) representing the image (710) from which the object (730) has been removed. To generate the background region (740), the electronic device may execute a generative artificial intelligence model or communicate with a server configured to execute the generative artificial intelligence model.
[0114] As described above with reference to FIGS. 3 to 5, the electronic device can use the location of the object (730) within the image (710) to determine or determine whether the object (730) is floating. In one embodiment of processing the exemplary image (710) of FIG. 7, the electronic device can determine that the object (730) is not floating because the object (730) includes the edge (714) of the image (710). If the object (730) is cut off from or overlaps the edge (714) of the image (710), the electronic device can determine that the object (730) is not floating.
[0115] As described above with reference to FIGS. 3 to 5, the electronic device can detect depth inversion of a background region (740) with respect to an object (730) using depth information (720). Referring to the exemplary depth information (720) of FIG. 7, it is assumed that a portion of a background region corresponding to another object (e.g., a bottle) different from the subject (e.g., a person) expressed through the object (730) has a shallower depth than the depth of the object (730). Based on this assumption, the electronic device can detect depth inversion. According to one embodiment, since the distance between the object (730) included in the image (710) and the other object and the camera is shorter than the distance between the object (730) and the camera, the depth value of the other object may be smaller than the depth value of the object (730).
[0116] An electronic device that detects depth inversion can apply a static stereoscopic visual effect to the image (710) as described above with reference to Table 3. If the static stereoscopic visual effect is not applied, other objects (e.g., bottles) positioned closer to the object (730) may move dynamically together with the background area, resulting in unnatural animation. Referring to Fig. 7, image frames (751, 752, 753) of a video generated by applying a stereoscopic visual effect to the image (710) are illustrated. The electronic device can generate the image frames (751, 752, 753) by reducing the deviation of the horizontal positions of the background area (640), as indicated by the information in Table 2, so that the background area (640) moves a relatively small distance (or at a slow speed) in the time domain. Referring to FIG. 7, since the background area (640) moves a relatively small distance, within the image frames (751, 752, 753) sequentially displayed within the time interval between t1 and t3, the object (730) may not overlap with a portion (719) representing a subject (e.g., a bottle) in the background area having a relatively close depth value. For example, the object (730) may not obscure the portion (719).
[0117] FIG. 8 illustrates the operation of an electronic device applying a stereoscopic visual effect to an exemplary image (810) according to some embodiments. The electronic device of FIG. 8 may include the electronic device (101) of FIG. 1 and / or FIG. 2 . The operation performed by the electronic device of FIG. 8 may be related to at least one of the operations of FIGS. 3 to 5 .
[0118] Referring to FIG. 8, an exemplary image (810) is illustrated. The electronic device can identify depth information (820) and / or an object (830) from the image (810). Similar to what was described above with reference to FIGS. 6 and 7, the depth information (820) of FIG. 8 can represent depth values corresponding to each pixel of the image (810) based on the density of dots. The electronic device can execute a generative artificial intelligence model to generate a background area (840) from which the object (830) has been removed.
[0119] Referring to FIG. 8, the object (830) may be spaced from the edge of the image (810). The electronic device may determine that the object (830) is floating by using the location of the object (830) within the image (810). For example, the electronic device may determine that the object (830) is floating if the ratio between the distance (h2) between the bottom of the image (810) and the object (830) and the height (h1) of the image (810) is greater than a threshold ratio (e.g., 9%). The electronic device may detect or identify, using the depth information (820), a portion of the background area (840) having a depth value less than the depth value of the object (830) (e.g., a portion adjacent to the bottom of the image (810). The electronic device, having identified both floating and depth inversion of the object (830), may apply a static visual effect to the image (810) based on Table 3.
[0120] Referring to FIG. 8, image frames (851, 852, 853) of a video are illustrated, which are generated by applying a static stereoscopic visual effect to an image (810). The electronic device can display the image frames (851, 852, 853) representing a background region (840) that moves a relatively small distance (or at a slow speed). Since the background region (840) moves a relatively small distance with respect to the object (830), the object (830) may not be separated from the shadow expressed in the background region (840) within the image frames (851, 852, 853) sequentially displayed within the time interval between t1 and t3. For example, the electronic device can apply the static stereoscopic visual effect to the image (810) so as not to express an unnatural motion, such as an object (830) separated from a shadow. For example, the electronic device may generate or display image frames (851, 852, 853) representing natural motion of an object (830) and a video including the image frames (851, 852, 853).
[0121] FIG. 9 illustrates the operation of an electronic device applying a stereoscopic visual effect to an exemplary image (910), according to one embodiment. The electronic device of FIG. 9 may include the electronic device (101) of FIG. 1 and / or FIG. 2 . The operation performed by the electronic device of FIG. 9 may be related to at least one of the operations of FIGS. 3 to 5 .
[0122] Referring to FIG. 9, an exemplary image (910) is illustrated. In response to an input for applying a three-dimensional visual effect to the image (910), the electronic device can obtain depth information (920) corresponding to the image (910). Referring to FIG. 9, a distribution of depth values included in the depth information (920) is illustrated based on the density of dots, similar to that described above with reference to FIGS. 6 to 8. The electronic device can perform object recognition on the image (910) to identify an object (930). The electronic device can generate or obtain a background area (940) representing the image (910) from which the object (930) has been removed.
[0123] Referring to FIG. 9, in an image (910) capturing a jumping person, an object (930) may be spaced apart from the edge of the image (910). Since the depth value of the object (930) in the image (910) is less than the depth value of the background area, depth inversion may not be detected. For example, when capturing a jumping person, depth inversion may not be detected because the probability that a subject closer than the person (e.g., a road between the person and the camera) will be excluded from the image (910) increases.
[0124] As described above, the electronic device that detects the floating object (930) and fails to detect depth inversion may apply a stereoscopic visual effect related to the jump action to the image (910) as described above with reference to Table 3. Referring to FIG. 9, image frames (951, 952, 953) of a video generated by applying a stereoscopic visual effect related to the jump action to the image (910) are illustrated. The electronic device may apply a stereoscopic visual effect set to emphasize the jump action to the image (910). The stereoscopic visual effect may increase the movement speed and / or movement distance of the background area (940) with respect to the object (930) more than other stereoscopic visual effects (e.g., the static stereoscopic visual effect of FIG. 8 and / or the stereoscopic visual effect of FIG. 6).
[0125] As described above, according to one embodiment, the electronic device can apply a stereoscopic visual effect to be applied to the image (910) based on the position and / or depth inversion of the object (930) within the image (910). As described above with reference to Table 3, the stereoscopic visual effects can be categorized into stereoscopic visual effects related to dynamic actions, static stereoscopic visual effects, and stereoscopic visual effects related to jump actions. The electronic device can change or adjust the stereoscopic visual effect to be applied to the image (910) by adjusting parameters (e.g., horizontal positions per key frame of the background area (940)) of information (e.g., information based on the exemplary JSON of Table 2) for defining a specific stereoscopic visual effect. For example, the position, movement direction, and / or movement speed of a layer corresponding to the background area (940) can be changed or determined based on the floating and / or depth inversion of the object (930).
[0126] Hereinafter, with reference to FIG. 10, an exemplary operation of an electronic device for applying a three-dimensional visual effect to an image in which multiple subjects are captured is described.
[0127] FIG. 10 is a flowchart of operations performed by an electronic device according to one embodiment. The electronic device of FIG. 10 may include the electronic device (101) of FIG. 1 and / or FIG. 2. The operations of FIG. 10 may be performed by the electronic device (101) and / or the processor (210) of FIG. 2. The operations of FIG. 10 may be performed based on the execution of programs illustrated in FIG. 2 (e.g., a salient object detector (241), a depth information estimator (242), a depth inversion detector (243), a floating object detector (244), a pixel unprojector (245), and / or a 3D renderer (246)). The operations of FIG. 10 may be related to operation (330) of FIG. 3.
[0128] The order in which the operations of FIG. 10 are performed may vary depending on the embodiment. For example, according to one embodiment, an electronic device may perform the operations of FIG. 10 in a different order than that shown in FIG. 10, or may perform at least two operations substantially simultaneously.
[0129] Referring to FIG. 10, in operation (1010), according to one embodiment, a processor of an electronic device may identify a subject to which a three-dimensional visual effect may be applied from an image. Operation (1010) may be performed by executing a primary object detector (241). The processor may execute a model for detecting subjects (e.g., an artificial intelligence model trained to perform object recognition as described above) to identify one or more subjects associated with the image. From the model, the processor may obtain information indicating the positions of each of the one or more subjects included in the image. The information may indicate bounding boxes representing each of the positions.
[0130] Referring to FIG. 10, in operation (1020), a processor of an electronic device according to one embodiment may determine whether multiple subjects have been identified from an image. If one subject has been identified from the image (1020—No), the processor may perform operation (1040). If multiple subjects have been identified from the image (1020—Yes), the processor may perform operation (1030).
[0131] Referring to FIG. 10, in operation (1040), according to one embodiment, the processor of the electronic device may determine an area within an image corresponding to a single subject as an area to which a stereoscopic visual effect will be applied. When detecting a subject from an image, the processor may determine a portion of the image associated with the subject as an area to which a stereoscopic visual effect will be applied.
[0132] Referring to FIG. 10, in operation 1030, a processor of an electronic device according to an embodiment may determine an area within an image corresponding to one of a plurality of objects as an area to which a stereoscopic visual effect will be applied. The processor may determine an area corresponding to a subject closest to the camera at the time the image was captured, among the plurality of objects, as the area. For example, the processor may execute a depth information estimator (242) to obtain depth information corresponding to the image. Using the depth information, the processor may obtain depth values of portions of the image corresponding to the plurality of objects, respectively. Among the depth values, the electronic device may determine a portion having a minimum depth value as an area to which the stereoscopic visual effect of operation 1030 will be applied. By recognizing a plurality of objects, the electronic device may identify an object of interest of the user and determine the identified object of interest as the object of operation 1030. When recognizing multiple subjects, the electronic device can determine the area to which the three-dimensional visual effect of operation (1030) will be applied by using photos stored in the electronic device, photos uploaded to a social network service (SNS), and / or other images viewed by the user.
[0133] In one embodiment, the processor may determine the area of operation (1030) based on the types (or classes, or categories) of the plurality of subjects. For example, the electronic device may determine an area classified as having captured a frontal view of a human face as the area of operation (1030), independently of the depth value. If the frontal view of a human face is not captured, the electronic device may obtain depth values of portions of the image corresponding to the plurality of subjects using the depth information, and determine the portion having the minimum depth value among the depth values as the area of operation (1030).
[0134] Referring to FIG. 10, in operation (1050), a processor of an electronic device according to an embodiment may apply a stereoscopic visual effect based on a location and / or depth within an image of an area. The processor may perform the operations of FIGS. 3 to 5 for the area of operation (1050), thereby selecting or determining a stereoscopic visual effect to be applied to the image. For example, the processor may select or determine a stereoscopic visual effect to be applied to the image based on a floating of the area of operation (1050) and / or a depth inversion detected based on the area of operation (1050), among the first stereoscopic visual effect described with reference to FIG. 6, the second stereoscopic visual effect described with reference to FIGS. 7 and / or 8, and the third stereoscopic visual effect described with reference to FIG. 9.
[0135] Hereinafter, with reference to FIG. 11a and / or FIG. 11b, stereoscopic visual effects that can be expressed in the format of Table 2 and conditions under which the stereoscopic visual effects are selected are exemplarily described.
[0136] FIGS. 11A and 11B illustrate exemplary operations of an electronic device for determining a stereoscopic visual effect to be applied to an image, according to some embodiments. The electronic device of FIGS. 11A and / or 11B may include the electronic device (101) of FIGS. 1 and / or 2. The operations of FIGS. 11A and / or 11B may be performed by programs illustrated in FIG. 2 (e.g., a salient object detector (241), a depth information estimator (242), a depth inversion detector (243), a floating object detector (244), a pixel unprojector (245), and / or a 3D renderer (246). The operations of the electronic device of FIGS. 11A and / or 11B may be related to at least one of the operations of FIGS. 3 to 5 and / or 10.
[0137] Referring to FIGS. 11A and 11B , exemplary stereoscopic visual effects applicable to an image (image (120) of FIG. 1 , image (230) of FIG. 2 , and / or image (315) of FIG. 3 ) by an electronic device, and exemplary conditions under which each of the stereoscopic visual effects is selected (e.g., first condition (1110) to seventh condition (1170)) are illustrated together with the exemplary images.
[0138] Referring to FIG. 11A, an electronic device that identifies an image satisfying a first condition (1110) may apply a stereoscopic visual effect referred to as Dolly Zoom L (left) and / or Dolly Zoom R (right) to the image. The first condition (1110) may be a condition in which an object of the image corresponds to a face (e.g., a human face) and has a size exceeding a threshold ratio within the image, in order to satisfy the first condition (1110). The electronic device may determine whether an image satisfying the first condition (1110) has been identified by using the size of the object within the image and / or the type (or class or category) of the foreground object corresponding to the object.
[0139] A stereoscopic visual effect referred to as Dolly Zoom L(left) and / or Dolly Zoom R(right), corresponding to the first condition (1110) of FIG. 11a, may be set to enlarge a background area with respect to an object. The left and right of Dolly Zoom L(left) and / or Dolly Zoom R(right) may indicate a movement direction of the enlarged background area. When applying a stereoscopic visual effect to an image satisfying the first condition (1110), the electronic device may determine the extent to which the background area is enlarged and / or the movement speed of the background area according to the floating and / or depth inversion of the object. For example, an electronic device that detects depth inversion from an image satisfying the first condition (1110) may reduce the extent to which the background area is enlarged and / or the movement speed of the background area, as described above with reference to FIG. 7.
[0140] Referring to FIG. 11a, an electronic device that identifies an image satisfying the second condition (1120) may apply a three-dimensional visual effect referred to as Top L2R (left to right) and / or Top R2L (right to left) to the image. The second condition (1120) may be a condition in which, in order to satisfy the second condition (1120), an object of the image is located at the top center among the nine equally divided portions of the image, and the object is spaced from the bottom of the image (e.g., floating).
[0141] The stereoscopic visual effect referred to as Top L2R (left to right) and / or Top R2L (right to left), corresponding to the second condition (1120) of FIG. 11a, can be set so that the background region moves horizontally (e.g., from left to right (L2R) and / or from right to left (R2L)) with respect to the object. The stereoscopic visual effect can be set so that the virtual camera adjacent to the top rotates (e.g., a change in the direction of the virtual camera, which is set by the "camera_dir_x" variable of Table 2). When applying the stereoscopic visual effect to an image that satisfies the second condition (1120), the electronic device can determine the degree and / or speed at which the background region moves, based on the depth inversion. For example, the electronic device that detects depth inversion from an image that satisfies the second condition (1120) can reduce the movement speed and / or movement distance of the background region.
[0142] Referring to FIG. 11a, an electronic device that identifies an image satisfying the third condition (1130) may apply a three-dimensional visual effect, referred to as a slide, to the image. The third condition (1130) may be a condition in which, in order to satisfy the third condition (1130), an object of the image is located at the upper left or upper right among the nine equally divided portions of the image, and the object is separated from the lower portion of the image.
[0143] The stereoscopic visual effect referred to as slide, corresponding to the third condition (1130) of FIG. 11a, may be set to move the background area less than the stereoscopic visual effect referred to as Top L2R (left to right) and / or Top R2L (right to left). For example, the key frame deviation of the horizontal position of the background area in Table 2 (e.g., “layer_position[0].x”) may be less than the key frame deviation of the horizontal position of the background area included in the information for defining the stereoscopic visual effect referred to as Top L2R (left to right) and / or Top R2L (right to left).
[0144] Referring to FIG. 11a, an electronic device that identifies an image satisfying the fourth condition (1140) may apply a three-dimensional visual effect referred to as TL2TR (top-left to top-right) and / or TR2TL (top-right to top-left) to the image. The fourth condition (1140) may be a condition in which an object of the image is located at the upper left or upper right among the nine equally divided portions of the image in order to satisfy the fourth condition (1140).
[0145] The stereoscopic visual effect referred to as TL2TR and / or TR2TL, corresponding to the fourth condition (1140) of FIG. 11a, may be set such that the background area moves in the upper horizontal direction (e.g., from the upper left to the upper right (TL2TR) and / or from the upper right to the upper left (TR2TL)) with respect to the object. The stereoscopic visual effect referred to as TL2TR and / or TR2TL may be set such that the background area moves less than the stereoscopic visual effect referred to as Top L2R and / or Top R2L.
[0146] Referring to FIG. 11a, an electronic device that identifies an image satisfying the fifth condition (1150) may apply a three-dimensional visual effect referred to as L2R (left to right) and / or R2L (right to left) to the image. The fifth condition (1150) may be a condition in which, in order to satisfy the fifth condition (1150), an object of the image is located in the central portion and the central lower portion among the portions of the image divided into nine equal parts, and the object corresponds to the upper body of a person (e.g., torso).
[0147] The stereoscopic visual effect referred to as L2R and / or R2L, corresponding to the fifth condition (1150) of FIG. 11a, may be set such that the background area moves horizontally (e.g., from left to right and / or from right to left) relative to the foreground direction. The stereoscopic visual effect referred to as L2R and / or R2L may be set such that the background area moves less than the stereoscopic visual effect referred to as Top L2R and / or Top R2L.
[0148] Referring to FIG. 11b, an electronic device that identifies an image satisfying the sixth condition (1160) may apply a three-dimensional visual effect referred to as BL2TR (bottom-left to top-right) and / or BR2TL (bottom-right to top-left) to the image. The sixth condition (1160) may be a condition in which, in order to satisfy the sixth condition (1160), an object of the image is located in the central portion and the central lower portion among the portions of the image divided into nine equal parts, and the object does not correspond to the upper body of a person (e.g., corresponds to the entire body of a person).
[0149] The stereoscopic visual effect referred to as BL2TR and / or BR2TL, corresponding to the sixth condition (1160) of FIG. 11b, may be set such that the background region moves diagonally with respect to the foreground direction (e.g., from bottom-left to top-right and / or from bottom-right to top-left). The stereoscopic visual effect referred to as BL2TR and / or BR2TL may be set such that the background region moves less than the stereoscopic visual effect referred to as Top L2R and / or Top R2L.
[0150] Referring to FIG. 11b, an electronic device that identifies an image satisfying the seventh condition (1170) may apply a three-dimensional visual effect referred to as TL2BR (top-left to bottom-right) and / or TR2BL (top-right to bottom-left) to the image. The seventh condition (1170) may be a condition in which an object of the image is located in the left center portion, the left lower portion, the right center portion, and / or the right lower portion among the nine equally divided portions of the image, in order to satisfy the seventh condition (1170).
[0151] Referring to FIG. 11b, the stereoscopic visual effect referred to as TL2BR and / or TR2BL corresponding to the seventh condition (1170) may be set such that the background area moves diagonally with respect to the foreground direction (e.g., from top-left to bottom-right and / or from top-right to bottom-left). The stereoscopic visual effect referred to as TL2BR and / or TR2BL may be set such that the background area moves less than the stereoscopic visual effect referred to as Top L2R and / or Top R2L.
[0152] Although the exemplary mapping of the seven conditions and stereoscopic visual effects has been described with reference to FIGS. 11A and / or 11B , the embodiments are not limited thereto, and the mapping of the conditions and stereoscopic visual effects may be implemented differently depending on the electronic device. When any one of the above conditions is satisfied, the electronic device may change the movement speed and / or movement distance of the background area depending on the floating and / or depth inversion of the object. For example, when detecting depth inversion, the electronic device may reduce the movement speed and / or movement distance of the background area.
[0153] In one embodiment where the electronic device corresponds to a HMD (e.g., the HMD device (293) of FIG. 2), the degree to which the background area is moved (e.g., the movement distance and / or movement speed) to apply a stereoscopic visual effect may be less than the degree to which the background area is moved by an electronic device other than the HMD to apply a stereoscopic visual effect. For example, in a stereoscopic execution environment with binocular disparity, an electronic device that is an HMD may apply a relatively small stereoscopic visual effect.
[0154] FIG. 12 illustrates a user interface (UI) displayed by an electronic device (101) to adjust a three-dimensional visual effect according to one embodiment. Referring to FIG. 12, an exemplary state of the electronic device of FIG. 1 is illustrated. The exemplary state of FIG. 12 may be related to the state (192) of FIG. 1.
[0155] Referring to FIG. 12, an exemplary state of an electronic device (101) displaying a first video (150), which is a result of applying a stereoscopic visual effect to an image (e.g., image (120) of FIG. 1), is illustrated. The electronic device (101) may receive an input for applying a stereoscopic visual effect to an image including an object and a background region. Based on the input, the electronic device (101) may determine one of a plurality of stereoscopic visual effects using a position of an object within the image, a depth of the object, and a depth of the background region. The plurality of stereoscopic visual effects may be preset. The object and background regions may be identified using segmentation information obtained from the image using a model trained to perform object recognition. The depth of the object and the depth of the background region may be identified using depth information obtained from the image using a model trained to determine depth values of pixels included in the image.
[0156] In one embodiment, the electronic device (101) may apply the determined stereoscopic visual effect to the image to generate a video corresponding to the image (e.g., the first video (150) of FIG. 12). For example, depending on whether the position of the object is spaced apart from the edge of the image, the electronic device may determine one of the plurality of stereoscopic visual effects. For example, depending on whether the depth of at least a portion of the background area is shallower than the depth of the object, the electronic device may determine one of the plurality of stereoscopic visual effects.
[0157] Referring to FIG. 12, an exemplary screen including a first video (150) representing the result of applying a stereoscopic visual effect to an image is illustrated. The electronic device (101) may display, on the display (110), visual objects for adjusting properties of the stereoscopic visual effect applied to the first video (150) (e.g., the direction of movement of a background area, a distance of movement, and / or a speed of movement). For example, the electronic device (101) may display visual objects (1210) for adjusting the direction of movement of a background area, an object, and / or a virtual camera. Arrows included in each of the visual objects (1210) may indicate different directions of movement. Based on an input for selecting one of the visual objects (1210), the electronic device may set the direction of movement of the background area, the object, and / or the virtual camera to the direction of the arrow of the visual object corresponding to the input.
[0158] Referring to FIG. 12, the electronic device (101) may display a visual object (1220) on the display (110) for controlling the distance by which a background area, an object, and / or a virtual camera is moved within a virtual space. The visual object (1220) may be configured to be movable along a straight line representing a selectable range, referred to as a slider. Based on an input for moving the visual object (1220), the electronic device may set the movement distance of the background area, the object, and / or the virtual camera to a distance corresponding to the input.
[0159] Referring to FIG. 12, the electronic device (101) may display a visual object (1230) on the display (110) for controlling the movement speed of a background area, an object, and / or a virtual camera. The visual object (1230) may be an indicator on a slider. The electronic device, which receives an input related to the visual object (1230), may set or control the movement speed of the background area, the object, and / or the virtual camera to a speed corresponding to the input.
[0160] Referring to FIG. 12, the electronic device (101) can display visual objects (1240, 1250) for sharing and / or storing the first video (150) on the display (110). The electronic device (101) that has received an input related to the visual object (1240) can display a menu for transmitting the first video (150) to an external electronic device such as a messenger, email, and / or television. The electronic device (101) that has received an input related to the visual object (1250) can store the first video (150) in the memory of the electronic device (101) (e.g., the memory (220) of FIG. 2). However, the embodiment is not limited thereto, and the electronic device (101) may further display a visual object (e.g., a button containing designated text such as “save as wallpaper”) for setting the first video (150) as the background screen and / or lock screen of the electronic device (101).
[0161] FIG. 13 illustrates exemplary operations of an electronic device for generating depth information (720) from an image (710), according to one embodiment. The electronic device of FIG. 13 may include the electronic device (101) of FIG. 1 and / or FIG. 2. The operations of FIG. 13 may be performed based on the execution of programs illustrated in FIG. 2 (e.g., a salient object detector (241), a depth information estimator (242), a depth inversion detector (243), a floating object detector (244), a pixel unprojector (245), and / or a 3D renderer (246). The operations of FIG. 13 may be related to operations of the electronic device for obtaining depth information (e.g., operation (320) of FIG. 3).
[0162] Referring to FIG. 13, an electronic device can obtain depth information (720) by using feature points (or key points) of an image (710), referred to as landmarks. The feature points on the image (710), indicated by X, can be identified or determined within the image (710) based on an algorithm for searching for feature points. By analyzing the feature points, the electronic device can obtain depth information (720) corresponding to the image (710).
[0163] According to one embodiment, an electronic device can generate or obtain depth information (720) using a mesh model (1310) representing a three-dimensional shape of a subject. For example, the electronic device can change the shape and / or posture of a mesh model (1310) (or 3D model) having a basic shape of a human according to the posture of the human represented by the image (710), thereby obtaining a mesh model (1310) having the posture of the human represented by the image (710). Using the mesh model (1310), the electronic device can generate or obtain depth information (720) (e.g., depth information expressing the depth of each body part of the human in detail).
[0164] Although one embodiment for applying a stereoscopic visual effect to an image (710) has been described, the embodiment is not limited thereto. The electronic device may perform similar operations for applications that perform three-dimensional reconstruction of an image (710) based on a two-dimensional color distribution, such as Neural Radiance Fields (NeRF) and / or 3D Gaussian Splatting.
[0165] As described above, according to one embodiment, when applying a three-dimensional visual effect to an image, the electronic device may select or determine the three-dimensional visual effect to be applied to the image by utilizing the positional relationship and / or depth relationship of objects in the image. For example, the position, movement distance, and / or movement direction of an object, a background area, and a virtual camera may be determined based on the positional relationship and / or depth relationship.
[0166] FIG. 14 is a block diagram of an electronic device (1401) within a network environment (1400) according to various embodiments. Referring to FIG. 14 , in the network environment (1400), the electronic device (1401) may communicate with the electronic device (1402) via a first network (1498) (e.g., a short-range wireless communication network), or may communicate with at least one of the electronic device (1404) or the server (1408) via a second network (1499) (e.g., a long-range wireless communication network). In one embodiment, the electronic device (1401) may communicate with the electronic device (1404) via the server (1408). According to one embodiment, the electronic device (1401) may include a processor (1420), a memory (1430), an input module (1450), an audio output module (1455), a display module (1460), an audio module (1470), a sensor module (1476), an interface (1477), a connection terminal (1478), a haptic module (1479), a camera module (1480), a power management module (1488), a battery (1489), a communication module (1490), a subscriber identification module (1496), or an antenna module (1497). In some embodiments, the electronic device (1401) may omit at least one of these components (e.g., the connection terminal (1478)), or may have one or more other components added. In some embodiments, some of these components (e.g., sensor module (1476), camera module (1480), or antenna module (1497)) may be integrated into a single component (e.g., display module (1460)).
[0167] The processor (1420) may, for example, execute software (e.g., a program (1440)) to control at least one other component (e.g., a hardware or software component) of the electronic device (1401) connected to the processor (1420) and perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (1420) may store commands or data received from other components (e.g., a sensor module (1476) or a communication module (1490)) in a volatile memory (1432), process the commands or data stored in the volatile memory (1432), and store result data in a non-volatile memory (1434). According to one embodiment, the processor (1420) may include a main processor (1421) (e.g., a central processing unit or an application processor) or an auxiliary processor (1423) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (1421). For example, when the electronic device (1401) includes the main processor (1421) and the auxiliary processor (1423), the auxiliary processor (1423) may be configured to use less power than the main processor (1421) or to be specialized for a given function. The auxiliary processor (1423) may be implemented separately from the main processor (1421) or as a part thereof.
[0168] The auxiliary processor (1423) may control at least a portion of functions or states associated with at least one component (e.g., the display module (1460), the sensor module (1476), or the communication module (1490)) of the electronic device (1401), for example, on behalf of the main processor (1421) while the main processor (1421) is in an inactive (e.g., sleep) state, or together with the main processor (1421) while the main processor (1421) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (1423) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (1480) or a communication module (1490)). In one embodiment, the auxiliary processor (1423) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (1401) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (1408)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.
[0169] The memory (1430) can store various data used by at least one component (e.g., the processor (1420) or the sensor module (1476)) of the electronic device (1401). The data can include, for example, software (e.g., the program (1440)) and input data or output data for commands related thereto. The memory (1430) can include volatile memory (1432) or non-volatile memory (1434).
[0170] The program (1440) may be stored as software in memory (1430) and may include, for example, an operating system (1442), middleware (1444), or an application (1446).
[0171] The input module (1450) can receive commands or data to be used in a component of the electronic device (1401) (e.g., a processor (1420)) from an external source (e.g., a user) of the electronic device (1401). The input module (1450) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0172] The audio output module (1455) can output audio signals to the outside of the electronic device (1401). The audio output module (1455) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0173] The display module (1460) can visually provide information to an external party (e.g., a user) of the electronic device (1401). The display module (1460) may include, for example, a display, a holographic device, or a projector, and a control circuit for controlling the device. In one embodiment, the display module (1460) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.
[0174] The audio module (1470) can convert sound into an electrical signal, or vice versa. According to one embodiment, the audio module (1470) can acquire sound through the input module (1450), output sound through the sound output module (1455), or an external electronic device (e.g., electronic device (1402)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (1401).
[0175] The sensor module (1476) can detect the operating status (e.g., power or temperature) of the electronic device (1401) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (1476) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0176] The interface (1477) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (1401) with an external electronic device (e.g., the electronic device (1402)). In one embodiment, the interface (1477) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0177] The connection terminal (1478) may include a connector through which the electronic device (1401) may be physically connected to an external electronic device (e.g., the electronic device (1402)). In one embodiment, the connection terminal (1478) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0178] The haptic module (1479) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. In one embodiment, the haptic module (1479) may include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
[0179] The camera module (1480) can capture still images and videos. In one embodiment, the camera module (1480) may include one or more lenses, image sensors, image signal processors, or flashes.
[0180] The power management module (1488) can manage the power supplied to the electronic device (1401). According to one embodiment, the power management module (1488) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).
[0181] A battery (1489) may power at least one component of the electronic device (1401). In one embodiment, the battery (1489) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0182] The communication module (1490) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (1401) and an external electronic device (e.g., electronic device (1402), electronic device (1404), or server (1408)), and the performance of communication through the established communication channel. The communication module (1490) may operate independently from the processor (1420) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (1490) may include a wireless communication module (1492) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (1494) (e.g., a local area network (LAN) communication module, or a power line communication module). Any of these communication modules may communicate with an external electronic device (1404) via a first network (1498) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (1499) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a local area network or a wide area network)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (1492) may use subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (1496) to verify or authenticate the electronic device (1401) within a communication network such as the first network (1498) or the second network (1499).
[0183] The wireless communication module (1492) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimizing terminal power and connecting multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency communications (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (1492) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (1492) may support various technologies for securing performance in high-frequency bands, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (1492) may support various requirements specified in the electronic device (1401), an external electronic device (e.g., the electronic device (1404)), or a network system (e.g., the second network (1499)). According to one embodiment, the wireless communication module (1492) may support a peak data rate (e.g., 20 Gbps or more) for eMBB implementation, a loss coverage (e.g., 164 dB or less) for mMTC implementation, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC implementation.
[0184] The antenna module (1497) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (1497) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (1497) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (1498) or the second network (1499), may be selected from the plurality of antennas by, for example, the communication module (1490). A signal or power may be transmitted or received between the communication module (1490) and the external electronic device via the selected at least one antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (1497).
[0185] According to various embodiments, the antenna module (1497) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high frequency band.
[0186] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
[0187] According to one embodiment, commands or data may be transmitted or received between the electronic device (1401) and an external electronic device (1404) via a server (1408) connected to a second network (1499). Each of the external electronic devices (1402 or 1404) may be the same or a different type of device as the electronic device (1401). According to one embodiment, all or part of the operations executed in the electronic device (1401) may be executed in one or more of the external electronic devices (1402, 1404, or 1408). For example, when the electronic device (1401) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (1401) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (1401). The electronic device (1401) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (1401) may provide an ultra-low latency service using, for example, distributed computing or mobile edge computing. In another embodiment, the external electronic device (1404) may include an Internet of Things (IoT) device. The server (1408) may be an intelligent server utilizing machine learning and / or a neural network.According to one embodiment, an external electronic device (1404) or server (1408) may be included within the second network (1499). The electronic device (1401) may be applied to intelligent services (e.g., smart homes, smart cities, smart cars, or healthcare) based on 5G communication technology and IoT-related technology.
[0188] In one embodiment, a method for applying a stereoscopic visual effect to an image may be utilized. In one embodiment, a method for determining or adjusting the intensity of the stereoscopic visual effect application may be required based on the content of the image (e.g., a foreground object and / or a background object). As described above, according to one embodiment, an electronic device may include a display, a memory including one or more storage media, storing instructions, and at least one processor including a processing circuit. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display an image on the display. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive an input for applying a stereoscopic visual effect to the image based on the displaying of the image on the display. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify, based on the input, segmentation information representing an object in the image. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate, on the display, a first video representing a background area within the image displaced by a first distance beyond the object as a result of applying the stereoscopic visual effect, based on identifying, from the segmentation information, the object including an edge of the image.The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate a second video, representing the background region within the image, on the display, as a result of applying the stereoscopic visual effect, based on identifying the object spaced from the edge of the image from the segmentation information, a second distance beyond the object that is less than the first distance. In one embodiment, the electronic device may apply the stereoscopic visual effect to an image. In one embodiment, the electronic device may determine or adjust an intensity at which the stereoscopic visual effect is applied, depending on content of the image (e.g., the object and / or the background region).
[0189] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to perform inpainting on the object to obtain an inpainted region that replaces the object. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain the background region, which includes the inpainted region and the remaining region of the image that is different from the object.
[0190] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify depth information corresponding to the image. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate the second video from among the first video or the second video based on identifying, from the depth information, at least a portion of the background area having a depth less than a depth between the object and a camera that captured the image.
[0191] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate the first video using the segmentation information based on identifying, from the depth information, the object having a depth less than the depth of the background area.
[0192] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify depth information indicated by sensor data of a time-of-flight (ToF) sensor or a light detection and ranging (LiDAR) sensor that was acquired along with the image.
[0193] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify depth information corresponding to the image using a model trained to output depth values respectively corresponding to pixels of the image.
[0194] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify information representing different stereoscopic visual effects. The information, defined for key frames of a video to which the stereoscopic visual effect is applied, may include a horizontal position of a first layer corresponding to the background area, a horizontal position of a second layer corresponding to the object, and a horizontal position of a virtual camera that is moved for rendering the first video or the second video within a virtual space including the first layer and the second layer.
[0195] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to change the horizontal position of the first layer indicated by the information such that a deviation of the horizontal positions of the first layer across the key frames is reduced based on identifying the object spaced from the edge of the image. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate a virtual space including the virtual camera, the second layer positioned sequentially from the virtual camera, and the first layer. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to move the virtual camera, the second layer, and the first layer within the virtual space according to the information to generate the second video.
[0196] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain segmentation information corresponding to the image using a model for object recognition.
[0197] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display, on the display, a visual object representing storage of the result, together with the result of applying the stereoscopic visual effect, including the first video or the second video. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to store, based on receiving an input for the visual object, the first video or the second video being displayed on the display as the result.
[0198] As described above, in one embodiment, a method of an electronic device including a display may be provided. The method may include an operation of displaying an image on the display. The method may include an operation of receiving an input for applying a stereoscopic visual effect to the image based on the displaying of the image on the display. The method may include an operation of identifying segmentation information representing an object in the image based on the input. The method may include an operation of generating, on the display, a first video representing a background region within the image displaced by a first distance beyond the object as a result of applying the stereoscopic visual effect, based on identifying, from the segmentation information, the object that includes an edge of the image. The method may include an operation of generating, on the display, a second video representing, on the display, the background region within the image displaced by a second distance less than the first distance beyond the object as a result of applying the stereoscopic visual effect, based on identifying, from the segmentation information, the object that is displaced by the edge of the image.
[0199] For example, the method may include performing inpainting on the object to obtain an inpainted region that replaces the object. The method may include obtaining a background region that includes the inpainted region and the remaining region of the image that is different from the object.
[0200] For example, the method may include an operation of identifying depth information corresponding to the image. The operation of generating the second video may include an operation of generating the second video from among the first video or the second video based on identifying, from the depth information, at least a portion of the background area having a shallower depth than the depth between the object and the camera that captured the image.
[0201] For example, the operation of generating the first video may include an operation of generating the first video using the segmentation information based on identifying, from the depth information, the object having a shallower depth than the depth of the background area.
[0202] For example, the operation of identifying the depth information may include an operation of identifying the depth information indicated by sensor data of a time-of-flight (ToF) sensor or a light detection and ranging (LiDAR) sensor acquired together with the image.
[0203] For example, the operation of identifying the depth information may include an operation of identifying the depth information corresponding to the image using a model trained to output depth values corresponding to each pixel of the image.
[0204] For example, the method may include an operation of identifying information representing different stereoscopic visual effects. The information, defined for key frames of a video to which the stereoscopic visual effect is applied, may include a horizontal position of a first layer corresponding to the background area, a horizontal position of a second layer corresponding to the object, and a horizontal position of a virtual camera that is moved for rendering the first video or the second video within a virtual space including the first layer and the second layer.
[0205] For example, the method may include an operation of changing the horizontal position of the first layer indicated by the information such that a deviation of the horizontal positions of the first layer across the key frames is reduced based on identifying the object spaced apart from the edge of the image. The method may include an operation of generating a virtual space including the virtual camera, the second layer positioned sequentially from the virtual camera, and the first layer. The method may include an operation of moving the virtual camera, the second layer, and the first layer within the virtual space according to the information to generate the second video.
[0206] For example, the identifying action may include an action of obtaining segmentation information corresponding to the image using a model for object recognition.
[0207] For example, the action of generating the second video may include an action of displaying a visual object on the display, the visual object representing the storage of the result, together with the result of applying the stereoscopic visual effect, including the first video or the second video. The method may include an action of storing the first video or the second video displayed on the display as the result, based on receiving an input for the visual object.
[0208] In one embodiment, as described above, a non-transitory computer-readable storage medium storing instructions may be provided. The instructions, when executed by an electronic device including a display, may cause the electronic device to receive an input for applying a stereoscopic visual effect to an image including an object and a background region. The instructions, when executed by the electronic device, may cause the electronic device to determine, based on the input, one of a plurality of designated stereoscopic visual effects using a position of an object within the image, a depth of the object, and a depth of the background region. The instructions, when executed by the electronic device, may cause the electronic device to apply the determined stereoscopic visual effect to the image and generate a video corresponding to the image. The instructions, when executed by the electronic device, may cause the electronic device to display the generated video on the display.
[0209] For example, the instructions, when executed by the electronic device, may cause the electronic device to identify the object and the background region using segmentation information obtained from the image using a model trained to perform object recognition.
[0210] For example, the instructions, when executed by the electronic device, may cause the electronic device to identify the depth of the object and the depth of the background area using depth information obtained from the image using a model trained to determine depth values of pixels.
[0211] For example, the instructions, when executed by the electronic device, may cause the electronic device to determine which of the plurality of designated stereoscopic visual effects to use depending on whether the location of the object is spaced from an edge of the image.
[0212] For example, the instructions, when executed by the electronic device, may cause the electronic device to determine one of the plurality of designated stereoscopic visual effects based on whether a depth of at least a portion of the background area is shallower than the depth of the object.
[0213] According to one embodiment, an electronic device as described above may include a display, a memory including one or more storage media for storing instructions, and at least one processor including a processing circuit. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive an input for applying a stereoscopic visual effect to an image including an object and a background region. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to determine, based on the input, one of a plurality of designated stereoscopic visual effects using a position of an object within the image, a depth of the object, and a depth of the background region. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to apply the determined stereoscopic visual effect to the image and generate a video corresponding to the image. The above instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display the generated video on the display.
[0214] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify the object and the background region using segmentation information obtained from the image using a model trained to perform object recognition.
[0215] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify the depth of the object and the depth of the background area using depth information obtained from the image using a model trained to determine depth values of pixels.
[0216] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to determine which of the plurality of designated stereoscopic visual effects is selected based on whether the location of the object is spaced from an edge of the image.
[0217] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to determine which of the plurality of designated stereoscopic visual effects is selected based on whether a depth of at least a portion of the background area is shallower than the depth of the object.
[0218] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.
[0219] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0220] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0221] Various embodiments of the present document may be implemented as software (e.g., a program (1440)) including one or more instructions stored in a storage medium (e.g., an internal memory (1436) or an external memory (1438)) readable by a machine (e.g., an electronic device (1401)). For example, a processor (e.g., a processor (1420)) of the machine (e.g., an electronic device (1401)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.
[0222] According to one embodiment, the method according to various embodiments disclosed in this document may be provided as a computer program product. The computer program product may be traded between sellers and buyers as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or may be provided through an application store (e.g., Play Store). TM ) or directly between two user devices (e.g., smart phones), online distribution (e.g., downloading or uploading). In the case of online distribution, at least a portion of the computer program product may be at least temporarily stored or temporarily created in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0223] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
[0224] As used herein, the term "if" will be understood to mean "when, upon," "in response to determining," or "in response to detecting," depending on the context. Similarly, "if it is determined to," or "if [the stated condition or event] is detected," will optionally be understood to mean "upon determining," or "in response to determining," "upon detecting [the stated condition or event]," or "in response to detecting [the stated condition or event]."
[0225] The devices described above may be implemented as hardware components, software components, and / or a combination of hardware components and software components. For example, the devices and components described in the embodiments may be implemented using one or more general-purpose computers or special-purpose computers, such as a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing instructions and responding to them. The processing device may execute an operating system (OS) and one or more software applications running on the operating system. The processing device may also access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing device is sometimes described as being used alone; however, one of ordinary skill in the art will recognize that the processing device may include multiple processing elements and / or multiple types of processing elements. For example, a processing unit may include multiple processors, or a processor and a controller. Other processing configurations, such as parallel processors, are also possible.
[0226] Software may include a computer program, code, instructions, or a combination of one or more of these, which may configure a processing device to perform a desired operation or may independently or collectively command the processing device. The software and / or data may be embodied in any type of machine, component, physical device, computer storage medium, or device for interpretation by the processing device or for providing instructions or data to the processing device. The software may also be distributed over networked computer systems and stored or executed in a distributed manner. The software and data may be stored on one or more computer-readable recording media.
[0227] The method according to the embodiment may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. In this case, the medium may be one that continuously stores a computer-executable program or one that temporarily stores it for execution or download. In addition, the medium may be various recording means or storage means in the form of a single or multiple hardware combinations, and is not limited to a medium directly connected to a computer system, but may also be distributed over a network. Examples of the medium may include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and those configured to store program commands, including ROM, RAM, and flash memory. In addition, examples of other media may include recording media or storage media managed by app stores that distribute applications, sites that supply or distribute various software, servers, etc.
[0228] Although the embodiments described above have been described by way of limited examples and drawings, those skilled in the art will appreciate that various modifications and variations can be made based on the above teachings. For example, appropriate results can still be achieved even if the described techniques are performed in a different order than described, and / or components of the described systems, structures, devices, circuits, etc. are combined or combined in a different manner than described, or are replaced or substituted with other components or equivalents.
[0229] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims described below.< / object>
Claims
1. In electronic devices, display; A memory including one or more storage media for storing instructions; and comprising at least one processor, comprising a processing circuit; The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Displaying an image on the above display; Based on displaying the image, receiving an input for applying a three-dimensional visual effect to the image; Based on the above input, segmentation information representing an object in the image is identified; Identifying an object from the above segmentation information; Based on the object containing the edge of the image: Generating a first video representing a background area within the image that is moved a first distance beyond the object on the display, thereby applying the three-dimensional visual effect, and Based on the object spaced from the edge of the image: To cause the stereoscopic visual effect to be applied by generating a second video representing the background area within the image moved by a second distance on the display, wherein the second distance is shorter than the first distance. Electronic devices.
2. In claim 1, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Performing inpainting on the above object to obtain an inpainting area that replaces the object; and causing the background area to be obtained, which includes the inpainting area and the remaining area of the image that is different from the object; Electronic devices.
3. In claim 1, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Identifying depth information corresponding to the above image; The action of generating the above second video is: generating the second video from among the first video or the second video based on identifying at least a portion of the background area having a shallower depth than the depth between the object and the camera that captured the image from the depth information; Electronic devices.
4. In claim 3, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Based on identifying the object having a shallower depth than the depth of the background area from the depth information, the segmentation information is used to generate the first video. Electronic devices.
5. In claim 3, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Causing to identify the depth information indicated by the sensor data of a ToF (time-of-flight) sensor or a LiDAR (light detection and ranging) sensor acquired together with the image, Electronic devices.
6. In claim 3, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Using a model trained to output depth values corresponding to each pixel of the image, the depth information corresponding to the image is identified, Electronic devices.
7. In claim 1, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: To identify information representing different stereoscopic visual effects, said information being defined for key frames of a video that are a result of applying said stereoscopic visual effect, said information being: Horizontal position of the first layer corresponding to the above background area; The horizontal position of the second layer corresponding to the above object; and A horizontal position of a virtual camera, which is moved for rendering the first video or the second video within a virtual space including the first layer and the second layer; including, Electronic devices.
8. In claim 7, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Based on identifying the object spaced from the edge of the image, changing the horizontal position of the first layer indicated by the information so that the deviation of the horizontal positions of the first layer across the key frames is reduced; Generating a virtual space including the virtual camera, the second layer positioned sequentially from the virtual camera, and the first layer; and In the virtual space, according to the information, causing the virtual camera, the second layer, and the first layer to move to generate the second video. Electronic devices.
9. In claim 1, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Inputting the image into a model for object recognition, and receiving the segmentation information as an output from the model, thereby causing the segmentation information corresponding to the image to be obtained. Electronic devices.
10. In claim 1, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: On the display, display a visual object indicating the storage of the result, together with the result of applying the stereoscopic visual effect, including the first video or the second video; and Based on receiving input for the visual object, causing the first video or the second video to be displayed on the display as the result to be stored, Electronic devices.
11. In a method of an electronic device including a display, An action of displaying an image on the above display; An action of receiving an input for applying a three-dimensional visual effect to the image based on displaying the image; An operation of identifying segmentation information representing an object of the image based on the above input; An action of identifying the object from the segmentation information; Based on the object containing the edge of the image: An operation of applying the three-dimensional visual effect by generating a first video representing a background area within the image that is moved a first distance beyond the object on the display, and Based on the object spaced from the edge of the image: An operation of applying the stereoscopic visual effect by generating a second video representing the background area within the image beyond the object on the display, wherein the second distance is shorter than the first distance. method.
12. In claim 11, An operation of performing inpainting on the above object to obtain an inpainting area that replaces the object; and Further comprising an operation of obtaining the background area, which includes the inpainting area and the remaining area of the image that is different from the object. method.
13. In claim 11, Further comprising an operation of identifying depth information corresponding to the image, The action of generating the second video is, An operation of generating the second video from among the first video or the second video based on identifying at least a portion of the background area having a shallower depth than the depth between the object and the camera that captured the image from the depth information, method.
14. In claim 13, the operation of generating the first video comprises: An operation of generating the first video using the segmentation information based on identifying the object having a shallower depth than the depth of the background area from the depth information, method.
15. In claim 13, the depth information is: Identified based on the sensor data of the ToF (time-of-flight) sensor or the LiDAR (light detection and ranging) sensor, which was acquired with the above image. method.
Citation Information
Patent Citations
Facilitating true three-dimensional virtual representation of real objects using dynamic three-dimensional shapes
EP3274966B1
Method and system for measuring a quality of three-dimensional image
KR1020130060522A
Remote control device for 3D video system
KR1020130117772A
Method and device for implementing stereo imaging
KR1020150021522A
Method and system for scene image modification
US20230410337A1