Electronic device, method, and non-transitory computer-readable storage medium for reconstructing video
The electronic device reconstructs 2D videos into 3D videos using motion analysis and neural networks, addressing the cost and complexity of traditional 3D video generation methods by enhancing depth perception.
Patent Information
- Application Number
- PCT/KR2025/007510
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-09-10
- Filing Date
- 2025-05-30
- Publication Date
- 2026-01-29
AI Technical Summary
Existing methods for generating 3D stereoscopic video require multiple cameras or sensors, which are costly and difficult to implement without them.
An electronic device with processing components reconstructs a 2D video into a 3D video by identifying and analyzing motion data of visual objects within frames using modules like object identification, image processing, and reconstruction, applying techniques such as optical flow and neural networks to generate depth values and parallax effects.
The method effectively converts 2D videos into 3D videos, providing enhanced depth perception without the need for multiple cameras or sensors, thus reducing costs and implementation complexity.
Smart Images

Figure KR2025007510_29012026_PF_FP_ABST
Abstract
Description
Electronic device, method, and non-transitory computer-readable storage medium for reconstructing video
[0001] The following descriptions relate to electronic devices, methods, and non-transitory computer-readable storage media for reconstructing video.
[0002] 3D stereoscopic video can be provided using two 2D videos of an object captured from different angles. The 2D videos can correspond to the viewer's left and right eyes, respectively. There can be parallax between the 2D videos. While watching the 2D videos, the viewer can perceive the video as a 3D stereoscopic video.
[0003] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above is applicable as prior art related to the present disclosure.
[0004] An electronic device is provided. The electronic device may include a memory comprising one or more storage media storing instructions. The electronic device may include at least one processor comprising processing circuitry. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive an input for generating a second video reconstructed from a first video. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain, from the first video, first motion data of a first visual object in the first video and second motion data of a second visual object in the first video based on the input. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate the second video representing the first visual object moving in front of the second visual object based on the first motion data being greater than the second motion data.
[0005] A method is provided. The method can be performed on an electronic device. The method can include receiving an input for generating a second video reconstructed from a first video. The method can include obtaining, from the first video, first motion data of a first visual object in the first video and second motion data of a second visual object in the first video based on the input. The method can include generating, based on the first motion data being greater than the second motion data, the second video representing the first visual object moving in front of the second visual object.
[0006] A non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium may store one or more programs. The one or more programs may include instructions that, when executed by an electronic device, cause the electronic device to receive an input for generating a second video reconstructed from a first video. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain, from the first video, first motion data of a first visual object in the first video and second motion data of a second visual object in the first video based on the input. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate, based on the first motion data being greater than the second motion data, the second video representing the first visual object moving in front of the second visual object.
[0007] Figure 1 illustrates an example of a user watching a video through a display of an electronic device.
[0008] Figure 2a is a simplified block diagram of an exemplary electronic device.
[0009] Figure 2b illustrates an example of components for reconstructing video of an electronic device.
[0010] Figure 3 illustrates an example of a two-dimensional video being reconstructed within an electronic device.
[0011] Figure 4 illustrates an example of the operations of an electronic device for reconstructing a video.
[0012] Figure 5 illustrates an example of the operations of an electronic device that reconstructs frames that make up a video.
[0013] Figure 6 illustrates an example of how frames constituting a video are reconstructed.
[0014] Figure 7 illustrates an example of an operation for changing the angle of view of a video.
[0015] Figure 8 illustrates an example of the operations of an electronic device for playing back reconstructed video.
[0016] FIG. 9 is a block diagram of an electronic device within a network environment according to various embodiments.
[0017] It will be understood that like reference numerals throughout the drawings refer to like parts, components, and structures.
[0018] The terms used in this disclosure are used only to describe specific embodiments and may not be intended to limit the scope of other embodiments. The singular expression may include plural expressions unless the context clearly indicates otherwise. Terms used herein, including technical or scientific terms, may have the same meaning as commonly understood by those of ordinary skill in the art described in this disclosure. Terms defined in general dictionaries among the terms used in this disclosure may be interpreted as having the same or similar meaning in the context of the relevant technology, and shall not be interpreted in an idealized or overly formal sense unless explicitly defined in this disclosure. In some cases, even if a term is defined in this disclosure, it cannot be interpreted to exclude embodiments of the present disclosure.
[0019] The various embodiments of the present disclosure described below illustrate a hardware-based approach as an example. However, since the various embodiments of the present disclosure include techniques utilizing both hardware and software, the various embodiments of the present disclosure do not exclude a software-based approach.
[0020] In the following description, terms referring to data (e.g., data, sensing data, information, luminescence information, identification information), terms referring to values, terms for operational states (e.g., operation, process), terms referring to objects (e.g., visual objects, executable objects), terms referring to network entities, terms referring to components of devices, etc. are examples for convenience of explanation. Therefore, the present disclosure is not limited to the terms described below, and other terms having equivalent technical meanings may be used. In addition, terms such as '... part', '... device', '... thing', '... body', etc. used below may mean at least one shape structure or a unit that processes a function.
[0021] In addition, in the present disclosure, expressions such as "more than" or "less than" may be used to determine whether a specific condition is satisfied or fulfilled, but this is merely a description for expressing an example and does not exclude descriptions such as "more than" or "less than." A condition described as "more than" may be replaced with "more than," a condition described as "less than" may be replaced with "less than," and a condition described as "more than and less than" may be replaced with "more than and less than." In addition, hereinafter, "A" to "B" mean at least one of elements from A (including A) to B (including B). hereinafter, "C" and / or "D" mean at least one of "C" or "D," that is, including {"C", "D", "C" and "D"}.
[0022] Figure 1 illustrates an example of a user watching a video through a display of an electronic device.
[0023] Referring to FIG. 1, in state (100), a user (102) can watch a video (120). For example, an electronic device (101) can play the video (120). For example, the electronic device (101) can include a display (110). The electronic device (101) can display the video (120) through the display (110).
[0024] The electronic device (101) may include memory. For example, the electronic device (101) may play a video (120) stored in the memory. As a non-limiting example, the electronic device (101) may receive data related to the video (120) from an external electronic device (e.g., a server) through the communication circuit of the electronic device (101). For example, the electronic device (101) may stream the video (120) using the data.
[0025] In one embodiment, the video (120) may include a two-dimensional video. For example, the two-dimensional video may not provide a three-dimensional effect to the user (102). For example, the three-dimensional effect may be referred to as a three-dimensional effect. As a non-limiting example, the two-dimensional video may provide a three-dimensional effect to the user (102). For example, the two-dimensional video may provide a lower level of three-dimensional effect to the user (102) than the three-dimensional effect of the three-dimensional video. For example, the user (102) may experience a higher three-dimensional effect when watching a three-dimensional video than when watching a two-dimensional video. For example, the user (102) may experience an enhanced user experience as the video (120) provides a higher level of three-dimensional effect.
[0026] In one embodiment, the video (120) may include a three-dimensional video. For example, the three-dimensional video may be described as a video (120) that can be expressed in three dimensions on the display (110). For example, the three-dimensional video may be provided using two two-dimensional videos. For example, the two two-dimensional videos may be represented as videos that capture the same environment with different fields of view (FOV) using a camera. For example, the user (102) may perceive the depth of visual objects within the video (120) by simultaneously viewing the two two-dimensional videos while wearing special glasses (e.g., glasses including red and blue lenses).
[0027] For example, a three-dimensional video can provide a three-dimensional effect to a user (102) by being displayed through a display (110) for displaying three-dimensional video. For example, a user (102) can perceive the depth of visual objects within a video (120) by watching two two-dimensional videos displayed through a display (110) for displaying three-dimensional video. For example, a display (110) for displaying three-dimensional video can be described as a display (110) manufactured for displaying three-dimensional video.
[0028] In one embodiment, a three-dimensional video can be provided using multiple two-dimensional videos (or images). For example, multiple cameras (or camera systems) can be used to acquire multiple two-dimensional videos (or images). For example, multiple cameras can capture multiple videos (or images) by capturing the same environment with different fields of view (FOV). For example, a three-dimensional video can provide a sense of depth by utilizing the parallax of multiple videos (or images) captured with different FOVs.
[0029] In one embodiment, a 3D (three-dimensional) scanner may be used to generate a three-dimensional video. For example, the 3D scanner may scan an external object using an optical signal (e.g., white light, laser) emitted from the 3D scanner and received by the 3D scanner. For example, the 3D scanner may acquire depth information of the external object by scanning the external object. For example, a three-dimensional video may be generated using the depth information of the external object acquired by the 3D scanner.
[0030] For example, in state (100), the user (102) may prefer a three-dimensional video over a two-dimensional video because the three-dimensional video provides a greater sense of depth than the two-dimensional video. For example, the user (102) may perceive that the quality of a two-dimensional video is higher than that of a three-dimensional video.
[0031] For example, generating a 3D video may require capturing an environment using multiple cameras or sensors used to acquire depth information about external objects. For example, capturing an environment using multiple cameras or sensors may be relatively expensive due to the need for multiple cameras or sensors. For example, when capturing an environment using multiple cameras or sensors, it may be difficult to utilize a 2D video captured without using multiple cameras or sensors.
[0032] For example, a method may be required to utilize a two-dimensional video captured without using multiple cameras or sensors. For example, a method for generating a three-dimensional video based on a two-dimensional video without using multiple cameras or sensors may be implemented within an electronic device (101).
[0033] For example, a method for reconstructing a two-dimensional video may be implemented within an electronic device (101). For example, the electronic device (101) may include components (or hardware components) for providing a method for reconstructing a two-dimensional video. The components are described and exemplified in more detail with reference to FIG. 2A.
[0034] Figure 2a is a simplified block diagram of an exemplary electronic device.
[0035] Referring to FIG. 2A, the electronic device (201) may be one of various forms of mobile devices, such as a laptop, smartphones of various form factors (e.g., bar-type smartphones, foldable-type smartphones, or rollable-type smartphones), tablets, cellular phones, and other similar computing devices. The electronic device (201) may be referred to as a user device, a multi-function device, or a portable device.
[0036] The electronic device (201) may include at least one processor (200), a display (210), and a memory (220). For example, the electronic device (201) may be an example of the electronic device (101) of FIG. 1. For example, the electronic device (201) may correspond to the electronic device (101). For example, the electronic device (201) may be an example of the electronic device (801) of FIG. 8. For example, the electronic device (201) may correspond to the electronic device (801). For example, the display (210) may be an example of the display (110) of FIG. 1. For example, the display (210) may correspond to the display (110).
[0037] At least one processor (200) may be configured to control the display (210) and the memory (220). At least one processor (200) may be configured to execute instructions stored in the memory (220) to cause the electronic device (201) to perform at least some of the operations illustrated in the descriptions of FIGS. 3 to 7.
[0038] According to one embodiment, at least one processor (200) may include a hardware component for processing data based on executing instructions. The hardware component for processing data may include, for example, a central processing unit (CPU) (e.g., including processing circuitry). For example, the hardware component for processing data may include a graphic processing unit (GPU) (e.g., including processing circuitry). For example, the hardware component for processing data may include a display processing unit (DPU) (e.g., including processing circuitry). For example, the hardware component for processing data may include a neural processing unit (NPU) (e.g., including processing circuitry). At least one processor (200) may include one or more cores. For example, at least one processor (200) may have a multi-core processor structure such as a dual core, a quad core, or a hexa core.
[0039] According to one embodiment, the display (210) may include hardware components of the electronic device (201) used to display images and / or videos. For example, the display (210) may include light-emitting elements and circuits (e.g., transistors) that control the light-emitting elements to emit light. For example, each of the light-emitting elements may include an organic light emitting diode (OLED) or a micro LED, but is not limited thereto. For example, the display (210) may include a liquid crystal display (LCD).
[0040] According to one embodiment, the memory (220) may include a hardware component for storing data and / or instructions input to and / or output from at least one processor (200). The memory (220) may include, for example, volatile memory such as random-access memory (RAM) and / or non-volatile memory such as read-only memory (ROM). The volatile memory may include, for example, at least one of dynamic RAM (DRAM), static RAM (SRAM), cache RAM, and pseudo SRAM (PSRAM). The non-volatile memory may include, for example, at least one of programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), flash memory, hard disk, compact disc, and embedded multimedia card (EMMC).
[0041] The electronic device (201) illustrated in the description of FIG. 2A can execute at least some of the operations illustrated in the description of FIGS. 3 to 7. The operations of the electronic device (201) illustrated in the description of FIGS. 3 to 7 can be performed, executed, caused, or controlled by at least one processor (200).
[0042] Figure 2b illustrates an example of components for reconstructing video of an electronic device.
[0043] Referring to FIG. 2B, the electronic device (201) may include an object identification module (250), an image processing module (260), and a reconstruction module (270). The electronic device (201) may reconstruct a two-dimensional video (e.g., video (300) of FIG. 3) into a three-dimensional video (or a two-dimensional video capable of providing a three-dimensional effect) using the object identification module (250), the image processing module (260), and the reconstruction module (270).
[0044] According to one embodiment, the object identification module (250) can be used to identify a visual object in an input video. The electronic device (201) (e.g., at least one processor (200)) can identify a visual object within a frame (or image) constituting a video using the object identification module (250). The object identification module (250) can be referred to as a detection manager. The electronic device (201) can obtain, calculate, or estimate movement data of the visual object by identifying the visual object using the object identification module (250). The electronic device (201) can identify location information of the visual object and the outline of the visual object using the object identification module (250). For example, the object identification module (250) can identify a visual object within a frame using a video processing (or digital image processing) technique. For example, the object identification module (250) can identify a visual object within a frame using an artificial intelligence model (e.g., a neural network). The identification of a visual object within a frame by an electronic device (201) will be described and illustrated in more detail with reference to FIG. 4.
[0045] According to one embodiment, the electronic device (201) may identify motion data of a visual object within a frame using the object identification module (250). For example, the electronic device (201) may obtain motion data of pixels corresponding to a visual object in the frame using the object identification module (250). As a non-limiting example, the electronic device (201) may obtain motion data of all pixels constituting a frame using the object identification module (250). The electronic device (201) may obtain or identify motion data between frames constituting a video using the object identification module (250). For example, the object identification module (250) may obtain or identify motion data of a visual object using a video processing (or digital image processing) technique. For example, the video processing technique may include optical flow and object tracking. For example, the object identification module (250) may obtain or identify motion data of a visual object using an artificial intelligence model (e.g., a neural network). The acquisition of movement data of visual objects within a frame by the electronic device (201) will be described and illustrated in more detail with reference to FIG. 4.
[0046] According to one embodiment, the image processing module (260) may be used to crop a visual object within a frame (or image) constituting an input video. The image processing module (260) may be referred to as an image processing manager. The electronic device (201) may provide a frame of the input video and output data of the object identification module (250) to the image processing module (260). The electronic device (201) may crop a visual object from a frame using the image processing module (260). The electronic device (201) may perform inpainting on a cropped area from a frame using the image processing module (260). The electronic device (201) may generate a background area of the reconstructed video using the image processing module (260). For example, the image processing module (260) may identify a visual object within a frame using a video processing (or digital image processing) technique. For example, the image processing module (260) can identify visual objects within a frame using an artificial intelligence model (e.g., a neural network). The operation of the electronic device (201) cropping a visual object within a frame and performing inpainting on the cropped area will be described and exemplified in more detail with reference to FIG. 5.
[0047] According to one embodiment, the reconstruction module (270) may be used to generate metadata of a reconstructed video based on the output data of the object identification module (250) and the output data of the image processing module (260). The reconstruction module (270) may be referred to as a 3D reconstruction manager. The electronic device (201) may use the reconstruction module (270) to generate or calculate a coordinate system based on the reconstructed background area. The electronic device (201) may use the reconstruction module (270) to calculate a depth value of a visual object. A method by which the electronic device (201) calculates a depth value of a visual object will be described and exemplified in more detail with reference to FIG. 4.
[0048] According to one embodiment, the electronic device (201) may store data of a background area of a reconstructed video as metadata of the reconstructed video using the reconstruction module (270). The electronic device (201) may store data of a visual object (e.g., depth value, location information) as metadata of the reconstructed video using the reconstruction module (270).
[0049] According to one embodiment, the electronic device (201) can generate a three-dimensional video (or a two-dimensional video capable of providing a three-dimensional effect) by performing the above-described operations on each of the frames constituting the two-dimensional video using the object identification module (250), the image processing module (260), and the reconstruction module (270).
[0050] Figure 3 illustrates an example of a two-dimensional video being reconstructed within an electronic device.
[0051] According to one embodiment, referring to FIG. 3, at least one processor (200) of the electronic device (201) may generate a video (350) by performing reconstruction of a video (300). For example, the video (350) may be represented as a video reconstructed from the video (300). For example, the at least one processor (200) may generate a three-dimensional video (e.g., the video (350)) by performing reconstruction of a two-dimensional video (e.g., the video (200)). For example, the reconstruction may be described as generating a video (or image) that provides a three-dimensional effect based on a video (or image) that does not provide a three-dimensional effect. For example, the reconstruction may include converting a video (or image) that does not provide a three-dimensional effect into a video (or image) that provides a three-dimensional effect.
[0052] In one embodiment, the video (300) may represent a two-dimensional video. For example, the video (300) may not include depth values (or depth information) of visual objects (e.g., visual object (301), visual object (302), visual object (303)) within the video (300). However, the present invention is not limited thereto. For example, the video (300) may have depth values of at least some of the visual objects within the video (300). For example, the video (300) may include relative depth information of one or more visual objects within the video (300). The relative depth information of the one or more visual objects may be in the form of metadata. For example, the relative depth information may include a distance from a camera that captured the video (300) to a real-world object corresponding to one or more visual objects.
[0053] According to one embodiment, the video (300) may be composed of frames (e.g., frame (310), frame (320), and frame (330)). For example, at least one processor (200) may display the video (300) through the display (210) by sequentially displaying the frames in time order through the display (210).
[0054] For example, at least one processor (200) can calculate or obtain a depth value (or depth information) of a visual object (e.g., visual object (301), visual object (302), visual object (303)) within the video (300). For example, the depth value of the visual object can be identified, calculated, or obtained based on the movement of the visual object. For example, the movement of the visual object can be referred to as local motion.
[0055] For example, motion parallax may occur based on the movement of visual objects (e.g., visual object (301), visual object (302), visual object (303)). Motion parallax can be explained as a three-dimensional sensation that occurs due to the experiential factors of the observer. For example, motion parallax can be referred to as a component of monocular three-dimensional sensation. For example, motion parallax can be explained as a phenomenon in which, while observing a moving object, the observer perceives an object that is close to the observer as moving relatively fast, and the observer perceives an object that is far from the observer as moving relatively slow. In other words, the three-dimensional sensation felt by the observer may vary depending on the movement of the object. For example, due to motion parallax, the observer may perceive a fast-moving object as moving in front of a slow-moving object.
[0056] For example, at least one processor (200) can calculate, output, or estimate a depth value of a visual object within a video (300) based on a motion parallax of the visual object (e.g., visual object (301), visual object (302), visual object (303)).
[0057] For example, the video (350) may include a three-dimensional video. For example, the video (350) may include a two-dimensional video that can provide a three-dimensional effect. For example, the video (350) may include a two-dimensional video that can exhibit a three-dimensional effect. For example, the video (350) may include a 2.5-dimensional video. A 2.5-dimensional video may be referred to as a video that can provide a three-dimensional effect by applying a three-dimensional effect to a two-dimensional video.
[0058] For example, video (350) may be represented as a video reconstructed from video (300). For example, visual objects (351), visual objects (352), and visual objects (353) within video (350) may have depth values (or depth information).
[0059] According to one embodiment, the depth value of a visual object (351) can be identified, calculated, or obtained based on the movement of the visual object (301) within the video (300). The depth value of a visual object (352) can be identified, calculated, or obtained based on the movement of the visual object (302) within the video (300). The depth value of a visual object (353) can be identified, calculated, or obtained based on the movement of the visual object (303) within the video (300).
[0060] For example, frame (310) may include visual objects (301), visual objects (302), and visual objects (303). For example, frame (320) may include visual objects (301), visual objects (302), and visual objects (303). For example, between frames (310) and (320), visual objects (301) may represent movement (321), visual objects (302) may represent movement (322), and visual objects (303) may represent movement (323).
[0061] For example, the motion (321) may be greater than the motion (322). For example, the video (350) may represent a visual object (351) moving ahead of a visual object (352) based on the motion (321) being greater than the motion (322).
[0062] For example, the motion (322) may be greater than the motion (323). For example, the video (350) may represent a visual object (352) moving ahead of the visual object (353) based on the motion (322) being greater than the motion (323).
[0063] For example, a frame (330) may include a visual object (301), a visual object (302), and a visual object (303). For example, between frames (320) and (330), the visual object (301) may represent movement (331), the visual object (302) may represent movement (332), and the visual object (303) may represent movement (333).
[0064] For example, the motion (331) may be greater than the motion (332). For example, the video (350) may represent a visual object (351) moving ahead of a visual object (352) based on the motion (331) being greater than the motion (332).
[0065] For example, the motion (332) may be greater than the motion (333). For example, the video (350) may represent a visual object (352) moving ahead of the visual object (353) based on the motion (332) being greater than the motion (333).
[0066] In one embodiment, a visual object having a motion less than a reference motion within the video (300) may correspond to a background area within the video (350). For example, between frames (310) and (320), the motion (323) of the visual object (303) may be less than the reference motion. For example, the video (350) may include the visual object (353) in the background area (or background image layer) based on the motion (323) of the visual object (303) being less than the reference motion.
[0067] For example, between frames (320) and (330), the movement (323) of the visual object (303) may be less than the reference movement. For example, the video (350) may include the visual object (353) in the background area (or background image layer) based on the movement (333) of the visual object (303) being less than the reference movement.
[0068] According to one embodiment, the video (350) can express a three-dimensional effect of a visual object by utilizing the motion parallax of the visual object (e.g., visual object (351), visual object (352), visual object (353)) within the video (350). For example, the video (350) can express a three-dimensional effect by varying the degree of movement of the visual object according to the depth value of the visual object. For example, the difference in the degree of movement of each visual object can be proportional to the difference in the depth value of each visual object.
[0069] Figure 4 illustrates an example of the operations of an electronic device for reconstructing a video.
[0070] Referring to FIG. 4, in operation 401, at least one processor (200) of the electronic device (201) may receive an input for reconstructing a first video. For example, the first video may include the video (300) of FIG. 3. For example, the first video may include a two-dimensional video. For example, the first video may not include depth values (or depth information) of visual objects within the first video. However, the present invention is not limited thereto. For example, the first video may have depth values of at least some of the visual objects within the first video.
[0071] In one embodiment, the input for reconstructing the first video may include selecting the first video stored in the memory (220). For example, the input may mean a part of the user's body (e.g., a finger) touching the display (210) and / or an input via an input device (e.g., a mouse, a keyboard), but is not limited thereto.
[0072] In operation 403, at least one processor (200) may determine whether reconstruction of the first video is possible based on the input received. For example, at least one processor (200) may determine whether conditions for reconstruction of the first video are satisfied based on the input received. For example, the conditions for reconstruction of the first video may include a first condition, a second condition, and a third condition.
[0073] In one embodiment, the first condition may include that the first video has a motion of the first video that is greater than a reference motion. For example, the motion of the first video may be represented by the global motion of the first video. For example, at least one processor (200) may identify whether the first video includes global motion data of the first video that is greater than the reference motion data.
[0074] For example, a first video (e.g., video (300)) may include global motion data. The global motion data may be described as data of movement (or motion) occurring in frames constituting the first video (e.g., frame (310), frame (320), frame (330) of FIG. 3). The global motion data may be generated when a camera that captured the first video shakes or moves. The global motion data may be referenced as motion data of frames generated based on the movement of the camera that captured the first video.
[0075] According to one embodiment, at least one processor (200) may acquire or produce global motion data of the first video using sensor information within the first video. For example, the sensor information may include information from a gyro sensor.
[0076] According to one embodiment, at least one processor (200) may obtain global motion data of a first video using video image processing technology.
[0077] According to one embodiment, at least one processor (200) may acquire or estimate global motion data of the first video using an artificial intelligence model (e.g., a neural network). For example, the artificial intelligence model may be utilized by a pre-trained artificial intelligence model based on a deep learning algorithm. For example, at least one processor (200) may acquire global motion vectors and / or optical flow data of the first video using a neural network. For example, the global motion data may include global motion vectors and optical flow data.
[0078] In one embodiment, the second condition may include at least one processor (200) identifying another video that includes a region (or visual object) corresponding to a region (or visual object) within the first video. For example, the region corresponding to the region within the first video may represent the same environment (or external object) in which the first video was filmed. For example, the region corresponding to the region within the first video may be represented as having been filmed from a different angle within the same environment. For example, the other video may be used in conjunction with the first video when calculating a depth value of the region (or visual object) within the first video.
[0079] For example, at least one processor (200) may calculate a depth value of the region by using image parallax of the first video and the other video. For example, image parallax may occur due to a discrepancy between the positions of visual objects in the first video and the positions of visual objects in the other video, depending on the positions (or optical axes) of the cameras that captured the first video and the positions (or optical axes) of the cameras that captured the other video.
[0080] For example, the image parallax of the first video and the other video may be required to be greater than the reference image parallax. For example, the reference image parallax may be described as a criterion by which at least one processor (200) can calculate a depth value of the region using the image parallax of the first video and the other video.
[0081] In one embodiment, the third condition may include that the movement of at least one of the visual objects within the first video is greater than the movement of a background area within the first video. For example, the movement of the at least one visual object within the first video is greater than the movement of a background area within the first video may include that as the at least one visual object within the first video moves, the at least one visual object leaves the frame of the first video.
[0082] For example, at least one processor (200) may identify whether conditions for reconstructing the first video are satisfied based on the reception of the input. For example, at least one processor (200) may execute operation 405 based on whether the first video satisfies the third condition and / or whether the first video satisfies the first condition or the second condition.
[0083] According to one embodiment, at least one processor (200) may crop at least one visual object within a frame constituting the first video based on whether the first video satisfies the third condition and / or whether the first video satisfies the first condition or the second condition. Cropping at least one visual object within a frame constituting the first video will be described and illustrated in more detail with reference to FIGS. 5 and 6.
[0084] For example, at least one processor (200) may execute operation 411 based on the first video not satisfying the third condition. For example, at least one processor (200) may execute operation 411 based on the first video not satisfying the first condition and the second condition.
[0085] In operation 405, at least one processor (200) may obtain first motion data of a first visual object in a first video and second motion data of a second visual object in the first video. For example, at least one processor (200) may obtain motion data of a visual object (e.g., a first visual object, a second visual object) in the first video. For example, at least one processor (200) may identify a visual object in the first video to obtain motion data of the visual object in the first video. Identifying the visual object in the first video may include identifying location information of the visual object in the first video. For example, at least one processor (200) may obtain motion data of the visual object in the first video by comparing frames constituting the first video.
[0086] According to one embodiment, at least one processor (200) may identify visual objects within the first video using video processing (or digital image processing) techniques. For example, the video processing algorithm may include a Haar cascade algorithm, a histogram of oriented gradient (HOG) algorithm, and a scale invariant feature transform (SIFT) algorithm.
[0087] As a non-limiting example, at least one processor (200) may use an artificial intelligence model (e.g., a neural network) to identify visual objects within the first video. For example, the artificial intelligence model may be pre-trained based on a deep learning algorithm. For example, the deep learning algorithm may include, but is not limited to, the You Only Look Once (YOLO) algorithm and the Single Shot Detector (SSD) algorithm.
[0088] For example, at least one processor (200) may identify a first visual object within a first video and a second visual object within the first video. For example, at least one processor (200) may acquire, calculate, or estimate first motion data of the identified first visual object and second motion data of the identified second visual object using an artificial intelligence model (e.g., a neural network).
[0089] According to one embodiment, at least one processor (200) may acquire, calculate, or estimate motion data of an identified visual object within a first video. For example, at least one processor (200) may utilize an artificial intelligence model (e.g., a neural network) to acquire, calculate, or estimate motion data of a visual object within the first video. For example, the artificial intelligence model may be pre-trained based on a deep learning algorithm. For example, the deep learning algorithm may include a simple online real-time tracking (SORT) algorithm, but is not limited thereto.
[0090] According to one embodiment, at least one processor (200) may obtain motion data of identified visual objects within a first video using video processing (or digital image processing) techniques.
[0091] In operation 407, at least one processor (200) may calculate a first depth value of a first visual object within a first video and a depth value of a second visual object within the first video. For example, at least one processor (200) may calculate, obtain, or estimate a first depth value of a first visual object within a frame constituting the first video and a depth value of a second visual object within a frame constituting the first video.
[0092] For example, at least one processor (200) may calculate, obtain, or estimate a first depth value based on first motion data of a first visual object. For example, at least one processor (200) may calculate, obtain, or estimate a second depth value based on second motion data of a second visual object. For example, the depth value of a visual object may decrease as the motion data of the visual object increases. For example, the smaller the depth value of a visual object, the farther the visual object may be located from a background area. For example, a viewer may perceive that the larger the depth value of a visual object, the farther the visual object is located from the viewer.
[0093] For example, at least one processor (200) can calculate, obtain, or estimate a first depth value of a first visual object in a first video and a second depth value of a second visual object in the first video by utilizing the property that the greater the motion data of a visual object, the smaller the depth value of the visual object.
[0094] According to one embodiment, at least one processor (200) may calculate, obtain, or estimate a depth value of a visual object (e.g., a first visual object, a second visual object) within a first video based on an image parallax of the visual object. For example, the first video may have global motion data of the first video that is greater than the reference motion data when the first condition is satisfied in operation 403.
[0095] For example, a visual object within a first video having global motion data of the first video greater than reference motion data may have an image parallax. For example, at least one processor (200) may obtain the image parallax of the visual object within the first video by comparing frames constituting the first video having global motion data of the first video greater than the reference motion data. For example, the larger the image parallax of the visual object within the first video, the smaller the depth value of the visual object may be. For example, the image parallax may mean a difference between a first frame and a second frame, which is obtained by projecting a first frame (or image) including a visual object within the first video onto a second frame (or frame) including a visual object within the first video.
[0096] For example, at least one processor (200) may use video image processing technology to calculate, acquire, or estimate a depth value of a visual object acquired based on image parallax. For example, at least one processor (200) may use an artificial intelligence model (e.g., a neural network) to calculate, acquire, or estimate a depth value of a visual object acquired based on image parallax. For example, the artificial intelligence model may include a pre-trained neural network based on a deep learning algorithm.
[0097] According to one embodiment, at least one processor (200) may, in operation 403, calculate, obtain, or estimate a depth value of a visual object in the first video based on the first video and another video including a visual object corresponding to the visual object in the first video, if the first video satisfies the second condition.
[0098] For example, at least one processor (200) may calculate, obtain, or estimate a depth value of a visual object within the first video by utilizing an image parallax between the first video and the other video. For example, the image parallax between the first video and the other video may refer to a difference between a frame within the first video and a frame within the other video, which is obtained by projecting a frame containing a visual object within the first video onto a frame containing a visual object within the other video.
[0099] For example, at least one processor (200) may use video image processing technology to calculate, obtain, or estimate a depth value of a visual object obtained based on an image parallax between the first video and the other video. For example, at least one processor (200) may use an artificial intelligence model (e.g., a neural network) to calculate, obtain, or estimate a depth value of a visual object obtained based on an image parallax between the first video and the other video. For example, the artificial intelligence model may include a pre-trained neural network based on a deep learning algorithm.
[0100] For example, at least one processor (200) may calculate, obtain, or estimate a depth value of a visual object in the first video based on motion data of the visual object in the first video and motion data of the visual object in the other video. For example, at least one processor (200) may calculate, obtain, or estimate a depth value of the visual object in the first video as an average of a first depth value estimated (or calculated, obtained) based on the motion data of the visual object in the first video and a second depth value estimated based on the motion data of the visual object in the other video.
[0101] For example, at least one processor (200) can calculate, obtain, or estimate a depth value of a visual object in a first video by adding a first depth value estimated (or calculated, obtained) based on motion data of a visual object in a first video and a second depth value estimated based on motion data of a visual object in another video.
[0102] According to one embodiment, at least one processor (200) may calculate, obtain, or estimate a depth value of a visual object in the first video based on a first depth value calculated (or obtained, estimated) based on motion data of a visual object (e.g., a first visual object, a second visual object) in the first video and a second depth value calculated based on image parallax of the visual object in the first video. For example, at least one processor (200) may calculate the depth value of the visual object in the first video as an average of the first depth value and the second depth value. For example, at least one processor (200) may calculate the depth value of the visual object in the first video by adding the first depth value and the second depth value.
[0103] According to one embodiment, at least one processor (200) may receive user input. For example, at least one processor (200) may set a depth value of a visual object within a first video based on the user input. For example, at least one processor (200) may change a depth value of a visual object within the first video based on the user input and store the depth value in the memory (220).
[0104] According to one embodiment, at least one processor (200) may be capable of calculating, obtaining, or estimating a depth value of a visual object within a frame (or each of the frames) constituting the first video.
[0105] According to one embodiment, at least one processor (200) may store depth values of visual objects within the first video within metadata corresponding to the first video.
[0106] According to one embodiment, at least one processor (200) may calculate a depth value of a visual object (e.g., a first visual object, a second visual object) within a first video based on motion parallax. For example, the greater the motion parallax of a visual object, the smaller the depth value of the visual object. For example, the greater the motion parallax of a visual object, the further away the visual object may be from a background area. For example, the smaller the depth value of a visual object, the further away the visual object may be from a background area.
[0107] In operation 409, at least one processor (200) may generate a second video. For example, at least one processor (200) may generate the second video based on the first motion data that is greater than the second motion data. For example, at least one processor (200) may generate the second video based on the first depth value that is less than the second depth value. For example, the second video may represent a first visual object moving in front of a second visual object.
[0108] For example, the second video may be described as a video reconstructed from the first video. For example, the second video may include video (350) of FIG. 3. For example, the second video may provide a sense of depth (or a stereoscopic effect). For example, the second video may include a three-dimensional video. For example, the second video may include a 2.5-dimensional video.
[0109] For example, at least one processor (200) can reconstruct a second video from the first video. For example, the second video can enhance the viewer's user experience by providing a sense of three-dimensionality. For example, by watching the second video, the viewer can experience a greater sense of immersion than when watching the first video.
[0110] In one embodiment, the electronic device (201) may include a wearable device. For example, the wearable device may include a head mounted device (HMD). For example, the second video may be displayed through a display of the HMD device.
[0111] At operation 411, at least one processor (200) may output a notification indicating that the first video cannot be reconstructed. For example, at least one processor (200) may output a notification indicating that the first video cannot be reconstructed based on identifying that the first video cannot be reconstructed in operation 403. For example, outputting a notification indicating that the first video cannot be reconstructed may include displaying a message (or visual object) indicating that the first video cannot be reconstructed through the display (210). For example, the message may be displayed on a pop-up window displayed through the display (210) based on identifying that the first video cannot be reconstructed.
[0112] For example, identifying that the first video cannot be reconstructed may include failing to satisfy the first condition of operation 403 or the second condition of operation 403. For example, identifying that the first video cannot be reconstructed may include failing to satisfy the third condition of operation 403.
[0113] According to one embodiment, at least one processor (200) may obtain third motion data representing the motion of a background area within the first video from the first video. For example, at least one processor (200) may output a notification indicating that reconstruction of the first video is not possible based on failure to identify at least one visual object within the first video having motion data greater than the third motion data.
[0114] Figure 5 illustrates an example of the operations of an electronic device that reconstructs frames that make up a video.
[0115] Referring to FIG. 5, the operations of FIG. 5 may be related to operations 405, 406, and 407 of FIG. 4.
[0116] In operation 501, at least one processor (200) may identify at least one visual object within a frame constituting the first video (e.g., frame (310), frame (320), frame (330) of FIG. 3). For example, at least one processor (200) may identify at least one visual object within a frame constituting the first video using an artificial intelligence model (e.g., a neural network). For example, the artificial intelligence model may include a pre-trained neural network based on a deep learning algorithm. For example, the deep learning algorithm may include a YOLO (you only look once) algorithm and a SSD (single shot detector) algorithm, but is not limited thereto.
[0117] For example, at least one processor (200) may identify at least one visual object within the first video using a video processing (or digital image processing) technique. For example, the video processing algorithm may include a Haar cascade algorithm, a histogram of oriented gradient (HOG) algorithm, and a scale invariant feature transform (SIFT) algorithm.
[0118] In operation 503, at least one processor (200) may identify motion data of at least one visual object within a frame constituting the first video. For example, operation 503 may correspond to at least a portion of operation 405 of FIG. 4. For example, at least one processor (200) may obtain motion data of at least one visual object within the first video by comparing frames constituting the first video.
[0119] For example, at least one processor (200) may acquire, calculate, or estimate motion data of at least one visual object within a frame constituting the first video using an artificial intelligence model (e.g., a neural network). For example, the artificial intelligence model may include a pre-trained neural network based on a deep learning algorithm. For example, the deep learning algorithm may include a simple online real-time tracking (SORT) algorithm. However, the present invention is not limited thereto. At least one processor (200) may store motion data of at least one visual object within a frame constituting the first video in a memory (220).
[0120] In operation 505, at least one processor (200) can identify whether the motion data of at least one visual object within a frame constituting the first video is greater than reference motion data. For example, the reference motion data can be described as a criterion for determining whether at least one visual object within a frame constituting the first video is included in a background area.
[0121] For example, at least one processor (200) may execute operation 507 based on motion data of at least one visual object within a frame constituting the first video that is greater than the reference motion data. For example, at least one processor (200) may execute operation 509 based on motion data of at least one visual object within a frame constituting the first video that is less than the reference motion data.
[0122] In operation 507, at least one processor (200) may obtain a first image layer (e.g., image layer (630), image layer (640) of FIG. 6) by cropping at least one visual object within a frame constituting the first video based on motion data of at least one visual object within a frame constituting the first video that is larger than reference motion data. For example, cropping at least one visual object within a frame constituting the first video may include performing cropping on an area corresponding to at least one visual object within a frame constituting the first video. For example, at least one processor (200) may crop at least one visual object within a frame constituting the first video using an image clipper algorithm.
[0123] According to one embodiment, at least one processor (200) may crop at least one visual object within each of the frames constituting the first video based on whether motion data of the at least one visual object within at least one frame among the frames constituting the first video is greater than reference motion data. For example, at least one processor (200) may utilize object tracking when cropping the at least one visual object within each of the frames constituting the first video. For example, at least one processor (200) may identify at least one visual object in the frames constituting the first video using object tracking.
[0124] According to one embodiment, at least one processor (200) may apply a filter effect and / or image quality processing technique to at least one visual object within a frame constituting the first video when performing a crop on the at least one visual object. For example, the filter effect may include blur processing, dim processing, and sharpening processing.
[0125] According to one embodiment, a first image layer may be obtained by cropping at least one visual object within a frame constituting the first video based on motion data of at least one visual object within a frame constituting the first video, the motion data being greater than reference motion data. For example, the cropped at least one visual object may include a visual object having motion within the visual object (e.g., a visual object corresponding to a fountain with moving water). For example, a visual object having motion within the visual object may be expressed through a display (210) as a change in the shape of the visual object.
[0126] In operation 509, at least one processor (200) may acquire a frame including at least one visual object within a frame constituting the first video as a second image layer (e.g., image layer (620) of FIG. 6) based on motion data of at least one visual object within a frame constituting the first video that is smaller than reference motion data. For example, the second image layer may include an image layer in which an area corresponding to at least one other visual object is cropped. For example, the at least one other visual object may be referenced as at least one other visual object within a frame constituting the first video that is larger than reference motion data.
[0127] According to one embodiment, the second image layer may be referenced as a background image of the second video. At least one processor (200) may perform inpainting on a cropped area within the second image layer (e.g., cropped area (621), cropped area (622) of FIG. 6). Inpainting may be described as a technique for restoring an area within a cropped (or deleted) image. For example, at least one processor (200) may perform inpainting on the cropped area using an artificial intelligence model (e.g., a neural network). At least one processor (200) may generate an image to be used for inpainting based on image information surrounding the cropped area using the artificial intelligence model. An operation of at least one processor (200) performing inpainting on the cropped area will be described and exemplified in more detail with reference to FIG. 6.
[0128] In one embodiment, the background region of the second video may include a second image layer. For example, the background region of the second video may include a single background image generated by concatenating second image layers obtained from each of the frames constituting the first video.
[0129] In one embodiment, the background area of the second video may include a background image generated by stitching together the common background of the first video and the other video in a panoramic form, if the first video satisfies the second condition in operation 403. For example, when stitching together the common background of the first video and the other video in a panoramic form, the same image quality processing may be applied.
[0130] According to one embodiment, at least one processor (200) can identify at least one visual object (e.g., a visual object corresponding to the ground, a visual object corresponding to the sky, a visual object corresponding to a building, a visual object corresponding to a mountain) within a background area. For example, at least one processor (200) can calculate a depth value of at least one visual object within the background area. For example, at least one processor (200) can enhance a three-dimensional effect within the background area by using the depth value of at least one visual object within the background area.
[0131] Figure 6 illustrates an example of how frames constituting a video are reconstructed.
[0132] Referring to FIG. 6, FIG. 6 may be related to the operations of FIG. 5. Frame (610) may be described as a frame constituting a first video. For example, image layer (630), image layer (640), and image layer (650) may represent frames constituting a second video.
[0133] According to one embodiment, frame (610) may be described as a frame of a first video. For example, frame (610) may include visual objects (601), visual objects (602), and visual objects (603). For example, frame (610) may correspond to frame (320) of FIG. 3.
[0134] For example, at least one processor (200) can obtain an image layer (620), an image layer (630), and an image layer (640) by executing operations 505, 507, and 509 of FIG. 5.
[0135] For example, at least one processor (200) may perform a crop on an area (621) corresponding to the visual object (601) within the frame (610) based on motion data of the visual object (601) that is larger than the reference motion data. For example, at least one processor (200) may store the cropped area (621) within the memory (220). For example, at least one processor (200) may obtain an image layer (640) by performing a crop on an area (621) corresponding to the visual object (601) within the frame (610). For example, at least one processor (200) may perform the crop using an image clipper algorithm.
[0136] For example, at least one processor (200) may perform cropping on an area (622) corresponding to a visual object (602) within a frame (610) based on motion data of the visual object (602) that is larger than reference motion data. For example, at least one processor (200) may store the cropped area (622) within a memory (220). For example, at least one processor (200) may obtain an image layer (630) by performing cropping on an area (622) corresponding to a visual object (602) within a frame (610).
[0137] For example, at least one processor (200) may refrain from cropping or skip cropping an area corresponding to a visual object (603) within a frame (610) based on motion data of the visual object (603) that is smaller than reference motion data. For example, at least one processor (200) may obtain a frame (610) in which cropping is performed on an area (621) and an area (622) as an image layer (620). For example, the frame (610) in which cropping is performed on an area (621) and an area (622) may include a visual object (603). For example, the image layer (620) may be represented as a background area of a second video.
[0138] For example, at least one processor (200) may obtain frames constituting the second video based on the image layer (620), the image layer (630), and the image layer (640). For example, at least one processor (200) may perform inpainting on a cropped area (621) within the image layer (620) and a cropped area (622) within the image layer (620). For example, inpainting may be described as a technique for restoring a damaged image area. For example, inpainting may be described as a technique for restoring an area within a cropped (or deleted) image.
[0139] According to one embodiment, at least one processor (200) may perform inpainting on a cropped area (e.g., cropped area (621) and cropped area (622)) using an artificial intelligence model (e.g., a neural network). For example, at least one processor (200) may perform inpainting on a cropped area using an image generated by the artificial intelligence model. The image generated by the artificial intelligence model may be generated based on image information surrounding the cropped area. The image generated by the artificial intelligence model may provide a similar feeling to the image surrounding the cropped area.
[0140] For example, at least one processor (200) can obtain an image layer (650) by performing inpainting on a cropped area (621) and a cropped area (622) from an image layer (620).
[0141] According to one embodiment, at least one processor (200) may utilize data related to the cropped region (e.g., cropped region (621), cropped region (622)) when performing inpainting on the cropped region within the frame (610). For example, the data related to the cropped region may include data of a surrounding region of the cropped region within the frame (610).
[0142] For example, data related to the cropped area may include data of another frame constituting the first video (e.g., frame (310), frame (330) of FIG. 3). For example, at least one processor (200) may perform inpainting on a cropped area (e.g., cropped area (621), cropped area (622)) within frame (610) by using data of an area corresponding to the cropped area within the other frame. For example, at least one processor (200) may perform inpainting on a cropped area (e.g., cropped area (621), cropped area (622)) within frame (610) by stitching an area corresponding to the cropped area within the other frame.
[0143] For example, at least one processor (200) may obtain frames constituting a second video based on the image layer (630), the image layer (640), and the image layer (650). For example, the frames constituting the second video may place the image layer (630) and the image layer (640) on (or in front of) the image layer (650) representing the background area.
[0144] For example, when displaying a second video through the display (210), at least one processor (200) may display a visual object (602) based on a depth value of a visual object (602) within an image layer (630), based on an image layer (650) representing a background area. For example, when displaying a second video through the display (210), at least one processor (200) may display a visual object (602) based on a depth value of a visual object (601) within an image layer (640), based on an image layer (650) representing a background area. For example, at least one processor (200) may represent a visual object (601) within an image layer (640) as being located farther away from the visual object (602) within the image layer (630), based on motion data of the visual object (601) that is greater than motion data of the visual object (602).
[0145] According to one embodiment, at least one processor (200) may store information of image layers (e.g., image layer (630), image layer (640), image layer (650)) as one frame in memory (220). As a non-limiting example, at least one processor (200) may store each of the image layers (e.g., image layer (630), image layer (640), image layer (650)) in memory (220).
[0146] Figure 7 illustrates an example of an operation for changing the angle of view of a video.
[0147] Referring to FIG. 7, at least one processor (200) may display a video via a display (210). For example, the video may include a video reconstructed from a two-dimensional video (e.g., video (350) of FIG. 3, the second video, video (710), video (720) of FIG. 4).
[0148] For example, at least one processor (200) may change the angle of view of the video. For example, changing the angle of view of the video may include changing the angle of an area within the video (or frame) displayed through the display (210) around a background image (or image) of the video. For example, changing the angle of view of the video may include rotating an area within the video displayed through the display (210) around a background image of the video.
[0149] For example, at least one processor (200) may change the angle of view of the video, thereby allowing the display (210) to display the video differently. For example, at least one processor (200) may display a visual object contained within the video differently. For example, at least one processor (200) may display a different aspect of a visual object within the video based on the angle of view of the video.
[0150] For example, the video (710) may be described as a video reconstructed from a two-dimensional video. For example, the video (710) may be represented as a field of view (711) on the display (210). For example, at least one processor (200) may represent the back of the visual object (351) and the back of the visual object (352) by displaying the video (710) through the display (210).
[0151] For example, the video (720) may be represented as a view angle (721) on the display (210). For example, the video (720) may be represented as a video in which the view angle (711) of the video (710) is changed to the view angle (721). For example, at least one processor (200) may represent the front of the visual object (351) and the front of the visual object (352) by displaying the video (710) through the display (210).
[0152] According to one embodiment, at least one processor (200) may, when displaying a video (710) through a display (210), represent a visual object (e.g., a visual object (351), a visual object (352)) within the video (710) as a field of view (711).
[0153] As a non-limiting example, at least one processor (200) may receive an input for changing the angle of view of the video (710) while displaying the video (710) through the display (210). For example, the input may mean a part of the user's body (e.g., a finger) touching the display (210) and / or an input through an input device (e.g., a mouse, a keyboard), but is not limited thereto.
[0154] For example, at least one processor (200) can change the angle of view of the video (710) from the angle of view (711) to the angle of view (721) based on the input. For example, at least one processor (200) can display the video (720) through the display (210) by changing the angle of view of the video (710) from the angle of view (711) to the angle of view (721). A user of the electronic device (201) can watch the video (720) while changing the angle of view. The electronic device (201) can provide an immersive experience to the user by changing the angle of view of the video (720) according to the user input. The electronic device (201) can enhance the user experience of the electronic device (201) by changing the angle of view of the video (720).
[0155] Although the angle of view of the reconstructed video in FIG. 7 is depicted as angle of view (711) and angle of view (721), this is merely exemplary. For example, at least one processor (200) may, based on the input, change the angle of view of the reconstructed video within an angle that can express the background area and visual objects through the display (210).
[0156] According to one embodiment, at least one processor (200) may generate a second video reconstructed from the first video by executing the operations of FIG. 4. For example, an angle representing the back of a background area of the second video may not be included in the field of view of the second video. For example, an angle representing an area without data within a frame constituting the first video may not be included in the field of view of the second video.
[0157] For example, the angle of view of the second video may change in response to the movement of the first video. For example, at least one processor (200) may change the angle of view of the second video in response to the movement of the first video while displaying the second video through the display (210). For example, the movement of the first video may be represented as the global motion of the first video. For example, the movement of the first video may be referenced as the movement of the camera that captured the first video.
[0158] For example, at least one processor (200) can provide a three-dimensional effect to the user by changing the angle of view while displaying the second video through the display (210). For example, the greater the three-dimensional effect the user feels, the more enhanced the user experience can be.
[0159] Figure 8 illustrates an example of the operations of an electronic device for playing back reconstructed video.
[0160] Referring to FIG. 8, in operation 801, at least one processor (200) may identify a first frame in the reconstructed video. The at least one processor (200) may identify data corresponding to the first frame stored in the memory (220). For example, the at least one processor (200) may identify data of a background area corresponding to the first frame, a depth value of at least one visual object within the first frame, position information of at least one visual object within the first frame, and movement data of at least one visual object within the first frame.
[0161] In operation 803, at least one processor (200) may move at least one visual object on a background area of the reconstructed video according to motion data of at least one object within the first frame. For example, at least one processor (200) may move at least one visual object in a coordinate system based on the background area according to the motion data. For example, at least one processor (200) may move at least one visual object in the coordinate system using a depth value of at least one visual object and position information of at least one visual object.
[0162] In operation 805, at least one processor (200) may change the angle of view of the reconstructed video. As a non-limiting example, at least one processor (200) may maintain the angle of view of the reconstructed video.
[0163] At operation 807, at least one processor (200) may identify a second frame subsequent to the first frame in the reconstructed video. The at least one processor (200) may perform operations 803 and 805 for the second frame. For example, if the first frame is the last frame of the reconstructed video, the at least one processor (200) may terminate playback of the reconstructed video.
[0164] FIG. 9 is a block diagram of an electronic device within a network environment according to various embodiments. The electronic device (901) may include the electronic device (201). For example, the electronic device (201) may correspond to the electronic device (901).
[0165] Referring to FIG. 9, in a network environment (900), an electronic device (901) may communicate with an electronic device (902) via a first network (998) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (904) or a server (908) via a second network (999) (e.g., a long-range wireless communication network). In one embodiment, the electronic device (901) may communicate with the electronic device (904) via the server (908). According to one embodiment, the electronic device (901) may include a processor (920), a memory (930), an input module (950), an audio output module (955), a display module (960), an audio module (970), a sensor module (976), an interface (977), a connection terminal (978), a haptic module (979), a camera module (980), a power management module (988), a battery (989), a communication module (990), a subscriber identification module (996), or an antenna module (997). In some embodiments, the electronic device (901) may omit at least one of these components (e.g., the connection terminal (978)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (976), the camera module (980), or the antenna module (997)) may be integrated into one component (e.g., the display module (960)).
[0166] The processor (920) may, for example, execute software (e.g., a program (940)) to control at least one other component (e.g., a hardware or software component) of the electronic device (901) connected to the processor (920) and perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (920) may store commands or data received from other components (e.g., a sensor module (976) or a communication module (990)) in a volatile memory (932), process the commands or data stored in the volatile memory (932), and store result data in a non-volatile memory (934). According to one embodiment, the processor (920) may include a main processor (921) (e.g., a central processing unit or an application processor) or an auxiliary processor (923) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (921). For example, when the electronic device (901) includes the main processor (921) and the auxiliary processor (923), the auxiliary processor (923) may be configured to use less power than the main processor (921) or to be specialized for a given function. The auxiliary processor (923) may be implemented separately from the main processor (921) or as a part thereof.
[0167] The auxiliary processor (923) may control at least a portion of functions or states associated with at least one component (e.g., a display module (960), a sensor module (976), or a communication module (990)) of the electronic device (901), for example, on behalf of the main processor (921) while the main processor (921) is in an inactive (e.g., sleep) state, or together with the main processor (921) while the main processor (921) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (923) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (980) or a communication module (990)). In one embodiment, the auxiliary processor (923) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (901) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (908)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.
[0168] The memory (930) can store various data used by at least one component (e.g., the processor (920) or the sensor module (976)) of the electronic device (901). The data can include, for example, software (e.g., the program (940)) and input data or output data for commands related thereto. The memory (930) can include a volatile memory (932) or a non-volatile memory (934).
[0169] The program (940) may be stored as software in the memory (930) and may include, for example, an operating system (942), middleware (944), or an application (946).
[0170] The input module (950) can receive commands or data to be used in a component of the electronic device (901) (e.g., a processor (920)) from an external source (e.g., a user) of the electronic device (901). The input module (950) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0171] The audio output module (955) can output audio signals to the outside of the electronic device (901). The audio output module (955) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0172] The display module (960) can visually provide information to an external party (e.g., a user) of the electronic device (901). The display module (960) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. In one embodiment, the display module (960) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.
[0173] The audio module (970) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (970) can acquire sound through the input module (950), output sound through the sound output module (955), or an external electronic device (e.g., electronic device (902)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (901).
[0174] The sensor module (976) can detect the operating status (e.g., power or temperature) of the electronic device (901) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (976) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0175] The interface (977) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (901) with an external electronic device (e.g., the electronic device (902)). In one embodiment, the interface (977) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0176] The connection terminal (978) may include a connector through which the electronic device (901) may be physically connected to an external electronic device (e.g., the electronic device (902)). In one embodiment, the connection terminal (978) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0177] The haptic module (979) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. According to one embodiment, the haptic module (979) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
[0178] The camera module (980) can capture still images and videos. According to one embodiment, the camera module (980) may include one or more lenses, image sensors, image signal processors, or flashes.
[0179] The power management module (988) can manage the power supplied to the electronic device (901). According to one embodiment, the power management module (988) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).
[0180] A battery (989) may power at least one component of the electronic device (901). In one embodiment, the battery (989) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0181] The communication module (990) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (901) and an external electronic device (e.g., electronic device (902), electronic device (904), or server (908)), and the performance of communication through the established communication channel. The communication module (990) may operate independently from the processor (920) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (990) may include a wireless communication module (992) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (994) (e.g., a local area network (LAN) communication module, or a power line communication module). Any of these communication modules may communicate with an external electronic device (904) via a first network (998) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (999) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (992) may use subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (996) to verify or authenticate the electronic device (901) within a communication network such as the first network (998) or the second network (999).
[0182] The wireless communication module (992) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency communications (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (992) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (992) may support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (992) may support various requirements specified in the electronic device (901), an external electronic device (e.g., the electronic device (904)), or a network system (e.g., the second network (999)). According to one embodiment, the wireless communication module (992) can support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.
[0183] The antenna module (997) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (997) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (997) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (998) or the second network (999), may be selected from the plurality of antennas, for example, by the communication module (990). A signal or power may be transmitted or received between the communication module (990) and the external electronic device via the at least one selected antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (997).
[0184] According to various embodiments, the antenna module (997) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high frequency band.
[0185] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
[0186] According to one embodiment, commands or data may be transmitted or received between the electronic device (901) and an external electronic device (904) via a server (908) connected to a second network (999). Each of the external electronic devices (902 or 904) may be the same or a different type of device as the electronic device (901). According to one embodiment, all or part of the operations executed in the electronic device (901) may be executed in one or more of the external electronic devices (902, 904, or 908). For example, when the electronic device (901) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (901) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (901). The electronic device (901) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (901) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (904) may include an Internet of Things (IoT) device. The server (908) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (904) or the server (908) may be included in the second network (999).The electronic device (901) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0187] An electronic device as described above may include a memory comprising one or more storage media storing instructions. The electronic device may include at least one processor comprising processing circuitry. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive an input for generating a second video reconstructed from a first video. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain, from the first video, first motion data of a first visual object in the first video and second motion data of a second visual object in the first video based on the input. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate the second video representing the first visual object moving in front of the second visual object based on the first motion data being greater than the second motion data.
[0188] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate the second video representing the second visual object as a background area based on the second motion data of the second visual object within the first video that is less than the reference motion data.
[0189] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain, based on the input, from the first video, first motion data of the first visual object within a frame constituting the first video and second motion data of the second visual object within the frame constituting the first video. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain, based on the first motion data of the first visual object within the frame, a first depth value of the first visual object within the frame. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain, based on the second motion data being less than the first motion data, a second depth value of the second visual object within the frame being greater than the first depth value. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate the second video representing the first visual object moving in front of the second visual object based on the first depth value being less than the second depth value.
[0190] In one embodiment, the electronic device may include a display. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display, through the display, the second video representing the first visual object moving in front of the second visual object. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive, while displaying the second video through the display, an input for changing a field of view of the second video, and, based on the input for changing the field of view of the second video, change the field of view of the second video displayed through the display.
[0191] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain a first image layer including the first visual object by cropping an area corresponding to the first visual object within a frame constituting the first video based on the first motion data that is greater than the reference motion data. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain a frame, as a second image layer, including the second visual object, the area corresponding to the first visual object being cropped based on the second motion data that is less than the reference motion data. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain a third image layer by performing inpainting on the area within the second image layer. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate the second video, based on the first image layer and the third image layer, representing the first visual object moving in front of the second visual object.
[0192] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain the third image layer by performing inpainting on the region within the second image layer using data associated with the region within the frame.
[0193] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain, from the first video, third motion data representing motion of a background area within the first video based on the input. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to output a notification indicating that the reconstruction of the first video is not possible based on a failure to identify at least one visual object within the first video having fourth motion data greater than the third motion data. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain the first motion data and the second motion data from the first video based on a failure to identify at least one visual object within the first video having fourth motion data greater than the third motion data. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate the second video representing the first visual object moving in front of the second visual object based on the first motion data being greater than the second motion data, based on identifying the at least one visual object in the first video having the fourth motion data being greater than the third motion data.
[0194] A method performed on an electronic device, as described above, may include receiving an input for generating a second video reconstructed from a first video. The method may include obtaining, from the first video, first motion data of a first visual object in the first video and second motion data of a second visual object in the first video based on the input. The method may include generating, based on the first motion data being greater than the second motion data, the second video representing the first visual object moving in front of the second visual object.
[0195] According to one embodiment, the method may include generating the second video representing the second visual object as a background area based on the second motion data of the second visual object within the first video that is smaller than the reference motion data.
[0196] According to one embodiment, the method may include an operation of obtaining, based on the input, from the first video, first motion data of the first visual object within a frame constituting the first video and second motion data of the second visual object within the frame constituting the first video. The method may include an operation of obtaining, based on the first motion data of the first visual object within the frame, a first depth value of the first visual object within the frame. The method may include an operation of obtaining, based on the second motion data that is less than the first motion data, a second depth value of the second visual object within the frame that is greater than the first depth value. The method may include an operation of generating, based on the first depth value that is less than the second depth value, the second video representing the first visual object moving in front of the second visual object.
[0197] According to one embodiment, the electronic device may include a display. The method may include an operation of displaying, through the display, the second video representing the first visual object moving in front of the second visual object. The method may include an operation of receiving, while displaying the second video through the display, an input for changing an angle of view of the second video, and changing the angle of view of the second video displayed through the display based on the input for changing the angle of view of the second video.
[0198] According to one embodiment, the method may include an operation of obtaining a first image layer including the first visual object by cropping an area corresponding to the first visual object within a frame constituting the first video based on the first motion data that is greater than the reference motion data. The method may include an operation of obtaining, as a second image layer, a frame including the second visual object and in which the area corresponding to the first visual object is cropped based on the second motion data that is less than the reference motion data. The method may include an operation of obtaining a third image layer by performing inpainting on the area within the second image layer. The method may include an operation of generating, based on the first image layer and the third image layer, the second video representing the first visual object moving in front of the second visual object.
[0199] According to one embodiment, the method may include obtaining the third image layer by performing inpainting on the region within the second image layer using data related to the region within the frame.
[0200] In one embodiment, the method may include an action of obtaining, from the first video based on the input, third motion data representing motion of a background area within the first video. The method may include an action of outputting a notification indicating that the reconstruction of the first video is not possible based on failing to identify at least one visual object within the first video having fourth motion data greater than the third motion data. The method may include an action of obtaining, from the first video, the first motion data and the second motion data based on identifying the at least one visual object within the first video having the fourth motion data greater than the third motion data. The method may include an action of generating, based on identifying the at least one visual object within the first video having the fourth motion data greater than the third motion data, the second video representing the first visual object moving in front of the second visual object based on the first motion data being greater than the second motion data.
[0201] In a computer-readable storage medium having one or more programs stored thereon, as described above, the one or more programs may include instructions that, when executed by an electronic device, cause the electronic device to receive an input for generating a second video reconstructed from a first video. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain, from the first video, first motion data of a first visual object in the first video and second motion data of a second visual object in the first video based on the input. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate, based on the first motion data being greater than the second motion data, the second video representing the first visual object moving in front of the second visual object.
[0202] According to one embodiment, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate the second video representing the second visual object as a background area based on the second motion data of the second visual object within the first video that is less than the reference motion data.
[0203] According to one embodiment, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain, from the first video, the first motion data of the first visual object within a frame constituting the first video and the second motion data of the second visual object within the frame constituting the first video, based on the input. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain, based on the first motion data of the first visual object within the frame, a first depth value of the first visual object within the frame. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain, based on the second motion data that is less than the first motion data, a second depth value of the second visual object within the frame that is greater than the first depth value. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate the second video representing the first visual object moving in front of the second visual object based on the first depth value being less than the second depth value.
[0204] In one embodiment, the electronic device may include a display. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to display, through the display, the second video representing the first visual object moving in front of the second visual object. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to receive, while displaying the second video through the display, an input for changing a field of view of the second video, and, based on the input for changing the field of view of the second video, change the field of view of the second video displayed through the display.
[0205] According to one embodiment, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain a first image layer including the first visual object by cropping an area corresponding to the first visual object within a frame constituting the first video based on the first motion data that is greater than the reference motion data. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain a frame, as a second image layer, including the second visual object and in which the area corresponding to the first visual object is cropped based on the second motion data that is less than the reference motion data. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain a third image layer by performing inpainting on the area within the second image layer. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate the second video, based on the first image layer and the third image layer, representing the first visual object moving in front of the second visual object.
[0206] According to one embodiment, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain the third image layer by performing inpainting on the region within the second image layer using data associated with the region within the frame.
[0207] According to one embodiment, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain, from the first video, third motion data representing motion of a background area within the first video based on a previous input. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to output a notification indicating that the reconstruction of the first video is not possible based on a failure to identify at least one visual object within the first video having fourth motion data greater than the previous third motion data. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain, from the first video, the first motion data and the second motion data based on identifying the at least one visual object within the first video having fourth motion data greater than the previous third motion data. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to generate the second video representing the first visual object moving in front of the second visual object based on the first motion data being greater than the second motion data, based on identifying the at least one visual object in the first video having the fourth motion data being greater than the third motion data.
[0208] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.
[0209] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0210] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0211] Various embodiments of the present document may be implemented as software (e.g., a program (940)) including one or more instructions stored in a storage medium (e.g., a memory (220), an internal memory (936), an external memory (938)) readable by a machine (e.g., an electronic device (201), an electronic device (901)). For example, a processor (e.g., at least one processor (200), a processor (920)) of the machine (e.g., an electronic device (201), an electronic device (901)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.
[0212] According to one embodiment, the method according to various embodiments disclosed in this document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) via an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0213] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
1. In electronic devices, A memory including one or more storage media for storing instructions; and At least one processor comprising a processing circuit, The above instructions, when individually or collectively executed by the at least one processor, Receive input for generating a second video reconstructed from a first video, Based on the input, from the first video, first motion data of a first visual object in the first video and second motion data of a second visual object in the first video are obtained, and To generate the second video representing the first visual object moving in front of the second visual object based on the first motion data that is greater than the second motion data, causing the above electronic device, Electronic devices.
2. In claim 1, The above instructions, when individually or collectively executed by the at least one processor, To generate the second video representing the second visual object as a background area based on the second motion data of the second visual object in the first video which is smaller than the reference motion data, causing the above electronic device, Electronic devices.
3. In claim 1, The above instructions, when individually or collectively executed by the at least one processor, Based on the input, from the first video, the first motion data of the first visual object within the frame constituting the first video and the second motion data of the second visual object within the frame constituting the first video are obtained, Based on the first motion data of the first visual object within the frame, a first depth value of the first visual object within the frame is obtained, Based on the second motion data that is smaller than the first motion data, a second depth value of the second visual object within the frame that is larger than the first depth value is obtained, and Generate the second video representing the first visual object moving in front of the second visual object based on the first depth value being smaller than the second depth value. causing the above electronic device, Electronic devices.
4. In claim 1, Including more displays, The above instructions, when individually or collectively executed by the at least one processor, Displaying the second video representing the first visual object moving in front of the second visual object through the display, and While displaying the above second video through the above display: Receive an input for changing the angle of view of the second video, and To change the angle of view of the second video displayed through the display based on the input for changing the angle of view of the second video, causing the above electronic device, Electronic devices.
5. In claim 1, The above instructions, when individually or collectively executed by the at least one processor, Obtaining a first image layer including the first visual object by performing a crop on an area corresponding to the first visual object within a frame constituting the first video based on the first motion data that is greater than the reference motion data, Based on the second motion data that is smaller than the reference motion data, the frame including the second visual object and the cropped area corresponding to the first visual object is acquired as a second image layer, By performing inpainting on the area within the second image layer, a third image layer is obtained, and To generate the second video, based on the first image layer and the third image layer, representing the first visual object moving in front of the second visual object; causing the above electronic device, Electronic devices.
6. In claim 5, The above instructions, when individually or collectively executed by the at least one processor, By performing inpainting on the area within the second image layer using data related to the area within the frame, the third image layer is obtained. causing the above electronic device, Electronic devices.
7. In claim 1, The above instructions, when individually or collectively executed by the at least one processor, Based on the above input, third motion data representing the motion of a background area within the first video is obtained from the first video, Outputting a notification indicating that the reconstruction of the first video cannot be performed based on failing to identify at least one visual object in the first video having fourth motion data greater than the third motion data, and Based on identifying at least one visual object within the first video having the fourth motion data greater than the third motion data: From the first video, the first motion data and the second motion data are obtained, and To generate the second video representing the first visual object moving in front of the second visual object based on the first motion data that is greater than the second motion data, causing the above electronic device, Electronic devices.
8. In a method performed in an electronic device, An operation for receiving input for generating a second video reconstructed from a first video, Based on the input, an operation of obtaining first motion data of a first visual object in the first video and second motion data of a second visual object in the first video from the first video, and An operation of generating a second video representing the first visual object moving in front of the second visual object based on the first motion data that is greater than the second motion data, method.
9. In claim 8, An operation of generating a second video showing the second visual object as a background based on the second motion data of the second visual object in the first video that is smaller than the reference motion data, method.
10. In claim 8, Based on the input, an operation of obtaining, from the first video, the first motion data of the first visual object within a frame constituting the first video and the second motion data of the second visual object within the frame constituting the first video; An operation of obtaining a first depth value of the first visual object within the frame based on the first motion data of the first visual object within the frame; An operation of obtaining a second depth value of the second visual object within the frame that is greater than the first depth value, based on the second motion data that is smaller than the first motion data, and An operation of generating a second video representing the first visual object moving in front of the second visual object based on the first depth value being less than the second depth value, method.
11. In claim 8, The electronic device comprises a display, An action of displaying, through the display, the second video representing the first visual object moving in front of the second visual object, and While displaying the above second video through the above display: An operation for receiving an input for changing the angle of view of the second video, and An operation of changing the angle of view of the second video displayed through the display based on the input for changing the angle of view of the second video, method.
12. In claim 8, An operation of obtaining a first image layer including the first visual object by performing a crop on an area corresponding to the first visual object within a frame constituting the first video based on the first motion data that is greater than the reference motion data; An operation of acquiring a frame as a second image layer, wherein the frame includes the second visual object and the area corresponding to the first visual object is cropped based on the second motion data that is smaller than the reference motion data; An operation of obtaining a third image layer by performing inpainting on the area within the second image layer, and An operation of generating a second video representing the first visual object moving in front of the second visual object based on the first image layer and the third image layer, method.
13. In claim 12, An operation of obtaining the third image layer by performing inpainting on the region within the second image layer using data related to the region within the frame, method.
14. In claim 8, Based on the above input, an operation of obtaining third motion data representing the motion of a background area within the first video from the first video; An action of outputting a notification indicating that the reconstruction of the first video cannot be performed based on not identifying at least one visual object in the first video having fourth motion data greater than the third motion data, and Based on identifying at least one visual object within the first video having the fourth motion data greater than the third motion data: An operation of obtaining the first motion data and the second motion data from the first video, and An operation of generating a second video representing the first visual object moving in front of the second visual object based on the first motion data that is greater than the second motion data, method.
15. In a non-transitory computer-readable storage medium storing one or more programs, the one or more programs, when executed by an electronic device, Receive input for generating a second video reconstructed from a first video, Based on the input, from the first video, first motion data of a first visual object in the first video and second motion data of a second visual object in the first video are obtained, and To generate the second video representing the first visual object moving in front of the second visual object based on the first motion data that is greater than the second motion data, comprising instructions causing the electronic device to operate; Non-transitory computer-readable storage medium.
Citation Information
Patent Citations
Electronic device for processing image and method for controlling thereof
KR1020160134428A
Lighting device for a motor vehicle and motor vehicle headlamp having such a lighting device
KR1020220074732A
High-rate composting method using compost composition of livestock sludge
KR1020250076278A
Cosmetic composition containing low molecular weight proteins of calendula, lavender and camellia
KR1020250159330A
Wearable device for executing application based on information obtained by tracking external object and method thereof
US20240143067A1