Electronic device for performing video inpainting and operation method thereof

The electronic device addresses noise and distortion in image inpainting by using depth information to align and inpaint occluded areas, ensuring natural restoration and generating high-quality stereo content.

WO2025154935A1PCT designated stage expired Publication Date: 2025-07-24SAMSUNG ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/018854
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-18
Filing Date
2024-11-26
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

Existing image inpainting technologies face challenges in naturally restoring areas occluded by objects due to noise and distortion, particularly in converting images or videos between viewpoints, and lack sufficient training data for stereo datasets.

Method used

An electronic device employs a method and model that utilizes depth information to align and inpaint occluded areas by extracting and aligning image and depth features across frames, incorporating an attention mechanism and depth-aware inpainting to maintain temporal and spatial consistency, and generates training data through disparity manipulation.

Benefits of technology

The solution effectively fills occluded areas with natural-looking pixels, enabling seamless conversion between viewpoints and generating stereo images/videos, improving the quality and consistency of inpainted content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024018854_24072025_PF_FP_ABST
    Figure KR2024018854_24072025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed is a method for inpainting a video by an electronic device. The method may comprise the steps of: extracting an input feature including an image feature and a depth feature on the basis of a first frame to be restored, a mask, and a depth frame; selecting a second frame which is a frame preceding the first frame and associated with the first frame; aligning the input feature of the first frame with an output feature of the second frame on the basis of the depth features of the first frame and the second frame; acquiring an output feature by modifying the image feature on the basis of the depth feature of the aligned input feature; and generating a restored frame on the basis of the output feature.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device for performing video inpainting and method of operation thereof

[0001] An electronic device and method of operating the same are provided for performing image inpainting, which restores an image by filling in pixels within the image.

[0002] AI-powered technologies are being applied to image editing functions provided by electronic devices. Image inpainting technology can restore an image by filling in pixels in areas that are damaged or empty within the image (holes). When an electronic device performs inpainting, noise such as blurriness and distortion can occur in the inpainted image depending on the characteristics of the hole areas. In this paper, we propose a method for naturally inpainting areas occluded by objects within an image.

[0003] According to one aspect of the present disclosure, a method for an electronic device to inpaint a video may be provided. The method may include extracting input features including image features and depth features based on a first frame to be reconstructed, a mask, and a depth frame. The method may include selecting a second frame that is a frame prior to the first frame and is associated with the first frame. The method may include aligning an input feature of the first frame with an output feature of the second frame based on the depth features of the first frame and the second frame. The method may include modifying an image feature based on the depth feature of the aligned input feature to obtain an output feature. The method may include generating a reconstructed frame based on the output feature.

[0004] According to one aspect of the present disclosure, an electronic device for performing inpainting can be provided. The electronic device can include a memory storing one or more instructions; and one or more processors executing the one or more instructions stored in the memory. The one or more processors can, by executing the one or more instructions, extract input features including image features and depth features based on a first frame to be reconstructed, a mask, and a depth frame. The one or more processors can, by executing the one or more instructions, select a second frame that is a frame prior to the first frame and associated with the first frame. The one or more processors, by executing the one or more instructions, can align an input feature of the first frame with an output feature of the second frame based on depth features of the first frame and the second frame. The one or more processors may obtain output features by modifying image features based on depth features of the aligned input features by executing the one or more instructions. The one or more processors may generate a reconstructed frame based on the output features by executing the one or more instructions.

[0005] According to one aspect of the present disclosure, a computer-readable recording medium having recorded thereon a program for performing inpainting provided by an electronic device, and executing any one of the methods described above and below may be provided.

[0006] FIG. 1 is a drawing for explaining an operation of an electronic device inpainting an image according to one embodiment of the present disclosure.

[0007] FIG. 2 is a flowchart illustrating an operation of an electronic device inpainting a video according to one embodiment of the present disclosure.

[0008] FIG. 3 is a diagram for explaining an image inpainting process performed by an electronic device according to one embodiment of the present disclosure.

[0009] FIG. 4A is a diagram illustrating an operation of an electronic device inpainting an image according to one embodiment of the present disclosure.

[0010] FIG. 4b is a diagram illustrating a detailed architecture of the inpainting model of the present disclosure.

[0011] FIG. 5 is a diagram for explaining an operation of an electronic device according to one embodiment of the present disclosure to align a current frame with reference to a previous frame.

[0012] FIG. 6 is a diagram for explaining an operation of an electronic device according to one embodiment of the present disclosure to align a current frame with reference to a previous frame.

[0013] FIG. 7 is a diagram for explaining an operation of an electronic device inpainting an image according to one embodiment of the present disclosure.

[0014] FIG. 8 is a diagram illustrating a depth feature generated by an electronic device according to one embodiment of the present disclosure.

[0015] FIG. 9 is a diagram for explaining an operation of an electronic device restoring pixel information according to an embodiment of the present disclosure.

[0016] FIG. 10 is a drawing for explaining an operation of an electronic device inpainting an image according to one embodiment of the present disclosure.

[0017] FIG. 11 is a diagram illustrating an operation of an electronic device generating training data according to an embodiment of the present disclosure.

[0018] FIG. 12 is a drawing for explaining an embodiment of an electronic device performing an inpainting operation according to an embodiment of the present disclosure.

[0019] FIG. 13 is a drawing for explaining an embodiment of an electronic device performing an inpainting operation according to an embodiment of the present disclosure.

[0020] FIG. 14 is a block diagram illustrating a configuration of an electronic device according to one embodiment of the present disclosure.

[0021] FIG. 15 is a block diagram illustrating the configuration of a server according to one embodiment of the present disclosure.

[0022] The terms used in this specification will be briefly explained, followed by a detailed description of the present disclosure. The terms used in this disclosure have been selected from widely used and common terms, taking into account the functions of the present disclosure. However, these terms may vary depending on the intentions of those skilled in the art, precedents, the emergence of new technologies, etc. Furthermore, in certain cases, terms may be arbitrarily selected by the applicant, in which case their meanings will be described in detail in the relevant description. Therefore, the terms used in this disclosure should not be defined simply as names of terms, but rather based on their meanings and the overall content of the present disclosure.

[0023] Singular expressions may include plural expressions unless the context clearly indicates otherwise. Terms used herein, including technical or scientific terms, have the same meaning as commonly understood by a person of ordinary skill in the art described herein. Furthermore, terms containing ordinal numbers, such as "first" or "second," used herein may be used to describe various components, but such components should not be limited by such terms. Such terms are used solely to distinguish one component from another.

[0024] When a part of the specification is said to "include" a component, unless otherwise specifically stated, this does not exclude other components but rather implies the inclusion of other components. Furthermore, terms such as "part" and "module" used in the specification refer to a unit that processes at least one function or operation, which may be implemented in hardware, software, or a combination of hardware and software.

[0025] Below, with reference to the attached drawings, embodiments of the present disclosure are described in detail so that those skilled in the art can easily practice the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In the drawings, portions irrelevant to the description have been omitted for clarity of explanation, and similar reference numerals have been used throughout the specification to designate similar parts.

[0026] The present disclosure will be described below with reference to the attached drawings.

[0027] FIG. 1 is a drawing for explaining an operation of an electronic device inpainting an image according to one embodiment of the present disclosure.

[0028] An electronic device according to one embodiment may be a device capable of playing images and / or videos. For example, the electronic device may be, but is not limited to, a smartphone, a tablet PC, a laptop PC, a desktop PC, a TV, or a head-mounted display (HMD).

[0029] In one embodiment, an electronic device can perform an image or video inpainting operation. Inpainting refers to the process of creating pixels and filling in the hole areas of an image, which contain no pixel information. When the electronic device restores the hole areas of an image, the electronic device can restore the image based on the surrounding pixels of the hole areas to achieve a natural restoration that matches the context of the image.

[0030] In one embodiment, the inpainting operation of an electronic device may be incorporated into various types of image processing operations. For example, the electronic device may utilize inpainting when synthesizing images or videos from new perspectives.

[0031] Referring to FIG. 1, the electronic device can convert an original image (110) into a new image (120). In the example of FIG. 1, the original image (110) is an image in which the camera (100) views a scene from a first viewpoint, and the new image (120) is an image in which the camera (100) views a scene from a second viewpoint.

[0032] In order for an electronic device to convert an original image (110) into a new image (120), it must estimate the depth of the original image (110) and move objects and the background within the original image (110) based on the depth information. As a result, an occlusion area that was covered by an object at the viewpoint of the original image (110) may exist in the new image (120). In other words, the occlusion area is a hole area where no pixel information exists. The electronic device can perform an inpainting operation to fill in pixels in the hole area, thereby making the new image (120) a complete image. At this time, since the hole area filled in during the inpainting process is an occlusion area indicating the back of the object, the electronic device can fill in the pixels by referring only to the information behind the object. To this end, the inpainting applied in the present disclosure utilizes depth information to enable natural inpainting with the surrounding context.

[0033] In one embodiment, the electronic device can convert a monoscopic image (or video) into a stereoscopic image (or video). For example, the electronic device can provide a user with 3D content composed of the original image (110) and the new image (120), with the original image (110) being a left-eye image and the new image (120) being a right-eye image. In this case, the electronic device can perform an inpainting operation to fill in the hole area resulting from the viewpoint transformation of the image when generating the new image (120).

[0034] The specific operations by which the electronic device performs the inpainting task will be described in more detail through the drawings and descriptions thereof described below.

[0035] FIG. 2 is a flowchart illustrating an operation of an electronic device inpainting a video according to one embodiment of the present disclosure.

[0036] In operation S210, the electronic device can extract input features including image features and depth features based on a current frame to be restored, a mask, and a depth frame.

[0037] In one embodiment, the electronic device may generate a current frame to be restored, a mask, and a depth frame. The current frame to be restored may refer to an image in which a specific area within the image is restored by an inpainting process performed by the electronic device. The current frame to be restored may be an image including a hole area without pixel information within the image. The mask may include binary information indicating the hole area, which is the area to be restored. The depth frame may include depth information regarding an object, background, etc. included in the frame to be restored. The operation of the electronic device generating the current frame to be restored, the mask, and the depth frame is further described with reference to FIG. 3.

[0038] In one embodiment, an electronic device can extract image features containing pixel information using a current frame to be restored and a mask. The electronic device can extract the image features using an image encoder. The image encoder can input a frame (image) and a mask, and output image features generated by reconstructing or transforming the input data. The image features may be, but are not limited to, a feature map in the form of a multidimensional array.

[0039] In one embodiment, an electronic device can extract a depth feature containing depth information using a current frame to be reconstructed, a mask, and a depth frame. The electronic device can extract the depth feature using a depth encoder. The depth encoder can input a frame (image), a mask, and a depth frame (depth image), and reconstruct and / or transform the input data to output a depth feature. The depth feature may be, but is not limited to, a feature map in the form of a multidimensional array.

[0040] The electronic device can concatenate image features and depth features to generate input features. In this case, the input features include image features and depth features.

[0041] In operation S220, the electronic device can select a previous frame associated with the current frame.

[0042] In one embodiment, the previous frame associated with the current frame may be a frame temporally related to the current frame. For example, the previous frame may be the current frame F t F, the frame immediately before t-1 It can be. Or, the previous frame is a frame F that is a certain number of frames prior to the current frame. t-n It could be.

[0043] In one embodiment, the electronic device may select a previous frame based on its similarity to the current frame. Since the previous frame is a frame that can be referenced for inpainting the current frame, the electronic device may select a previous frame whose similarity to the current frame is greater than a preset value as the frame associated with the current frame.

[0044] In operation S230, the electronic device can align the input features of the current frame with the output features of the previous frame based on the depth features of the previous frame and the current frame.

[0045] In one embodiment, the electronic device can process input features for each frame of the video to obtain output features including inpainted pixel information. In other words, the output features of the previous frame may be generated as a result of inpainting the previous frame. The output features of the previous frame may include image features and depth features. The electronic device can use the output features corresponding to the previous frame in the inpainting process of the current frame to ensure that the current frame is naturally connected to the previous frames.

[0046] The operation of the electronic device to align the current frame with reference to the previous frame is further described with reference to FIGS. 5 and 6.

[0047] In operation S240, the electronic device can obtain an output feature by modifying an image feature based on a depth feature of the aligned input feature.

[0048] In one embodiment, the aligned input features may include depth features and image features. In this case, the depth features may include depth information that restores a hole region corresponding to a mask within a depth frame. In other words, since the depth features include depth information that restores a hole region, inputting the depth features to a depth decoder to restore a depth frame may generate a depth frame in which the hole region is filled.

[0049] An electronic device can modify an image feature using a depth feature including restored depth information. The electronic device can calculate a similarity between the depth information of a specific pixel and the depth information of surrounding pixels, based on the depth feature of the aligned input feature, and calculate and modify the value of the image feature based on the similarity. In this case, the output feature can include a depth feature including restored depth information and an image feature including restored pixel information.

[0050] The operation of obtaining output features by modifying image features to include restored pixel information by an electronic device is further described with reference to FIGS. 7 to 9.

[0051] In operation S250, the electronic device can generate a restored frame based on the output feature.

[0052] In one embodiment, the image features included in the output feature may include pixel information that restores the hole region. In this case, by inputting the image features to an image decoder to restore an image frame, an image frame in which the hole region is filled can be obtained.

[0053] FIG. 3 is a diagram for explaining an image inpainting process performed by an electronic device according to one embodiment of the present disclosure.

[0054] In one embodiment, an electronic device may obtain an original image (310). The original image (310) may be received by the electronic device from an external device (e.g., a server) or may be captured using a camera module included in the electronic device. The electronic device may perform a series of operations on the original image (310) to generate an inpainted image (350). The inpainted image (350) may be an image obtained by converting the viewpoint of the original image (310). For example, if the original image (310) is an image of a first viewpoint (e.g., a left viewpoint), the inpainted image (350) may be an image of a second viewpoint (e.g., a right viewpoint) different from the first viewpoint.

[0055] An electronic device can estimate the depth of an original image (310) and obtain a depth image (320). The electronic device can perform a depth estimation task using various known algorithms and / or a deep neural network for depth estimation tasks and obtain a depth image (320). The depth image (320) can include depth and position information of an object. For example, a pixel value can represent depth information, and a higher pixel value can represent an object that is farther away from the camera, i.e., a greater distance, and a lower pixel value can represent a closer distance.

[0056] An electronic device can generate a disparity map (330). The disparity map (330) may include disparity information indicating a visual difference between two images. For example, the value of each pixel in the disparity map (330) may indicate the disparity from that pixel to the corresponding pixel in another image. In this case, a larger disparity indicates that the object is closer in the two images, and a smaller disparity indicates that the object is further away in the two images.

[0057] An electronic device may generate a disparity map (330) to be used to convert the viewpoint of the original image (310) using the depth information of the depth image (320). The disparity map (330) may include information about the parallax between the original image (310) and another image, i.e., information about the direction and how much an object and background should move in the original image (310) in order to convert the viewpoint of the original image (310). For example, if the original image (310) is an image of a left viewpoint and an image of a right viewpoint is to be generated, the object in the original image (310) should move to the left, and the background of the original image (310) should move to the right. In addition, since the object is closer than the background in both images, the distance by which the object moves to the left should be greater than the distance by which the background moves to the right. In this case, the disparity map (330) may include information indicating the movement direction and distance of the object and background within the image.

[0058] An electronic device can generate a warped image (340) using an original image (310) and a disparity map (330). The warped image (340) refers to an image in which the shape, size, direction, etc. of the original image (310) or an object is warped based on the disparity map (330).

[0059] In one embodiment, the converted image (340) may include an occlusion area due to an object or the like within the original image (310) as the viewpoint of the original image (310) is converted. The occlusion area is an area without pixel information and may be referred to as a hole area. Since the converted image (340) includes an area without pixel information, it may also be referred to as an image to be restored. The electronic device may generate a mask including information indicating the hole area.

[0060] An electronic device can generate an inpainted image (350). The electronic device can perform an inpainting operation to fill pixel information in a hole area in a converted image (340) where there is no pixel information. The electronic device can utilize depth information when inpainting the converted image (340). In the example of FIG. 3, since the hole area (occluded area) of the converted image (340) is an area behind an object, when restoring the converted image (340), the pixel information corresponding to the hole area must be filled in with reference to the background area. With reference to the drawings below, operations for the electronic device to generate an inpainted image (350) based on an image to be restored and depth information will be described.

[0061] In one embodiment, the electronic device can convert a 2D image into a 3D image. Alternatively, the electronic device can convert a 2D video into a 3D video. For example, as illustrated in FIG. 3 , the electronic device can generate an inpainted image (350) of a second viewpoint (e.g., a right viewpoint) based on an original image (310) of a first viewpoint (e.g., a left viewpoint). The electronic device can provide the original image (310) and the inpainted image (350) together to provide a 3D image and / or a 3D video.

[0062] FIG. 4A is a diagram illustrating an operation of an electronic device inpainting an image according to one embodiment of the present disclosure.

[0063] Referring to FIG. 4A, the inpainting model (400) may include at least an image encoder (410), a depth encoder (420), an alignment module (430), an inpainting module (440), and an image decoder (450). The components of the inpainting model (400) may be composed of a series of codes and a combination of data to implement specific functions. Each component may process operations included in the task of inpainting an image and / or a video by an electronic device. Each component may interact with each other by transmitting data. The processor of the electronic device may process tasks corresponding to each module of the inpainting model (400) by executing the execution code of the inpainting model (400). Hereinafter, the components of the inpainting model (400) will be described.

[0064] The image encoder (410) can extract image features based on an input frame and a mask. The input frame may be an image that includes a hole region without pixel information, and may be referred to as a frame to be restored. The mask may include information about the hole region.

[0065] The depth encoder (420) can extract a depth feature based on an input frame, a mask, and a depth frame. The depth feature may include depth information of the input frame, but may also include depth information that restores a hole area corresponding to the mask. In other words, the depth encoder can extract a depth feature based on the input frame, the mask, and the depth frame, and generate feature values ​​such that the depth feature also includes depth information corresponding to the hole area. In this case, when the depth feature is input to the depth decoder to restore the depth frame, a depth frame in which the hole area is filled can be generated. A detailed description of the depth feature is further described with reference to FIG. 8. The image feature and the depth feature can be passed to the alignment module (430).

[0066] The alignment module (430) can align features (image features and depth features) of the previous frame with features (image features and depth features) of the current frame to maintain temporal consistency in video inpainting. The alignment module (430) can calculate new aligned feature values, for example, using an attention mechanism.

[0067] The inpainting module (440) can restore pixel information by referring to the background area while considering depth information when generating pixel information of an occluded area behind an object. The inpainting module (430) can calculate new inpainted feature values, for example, using an attention mechanism.

[0068] The image decoder (450) can transform the output feature to generate an output frame. The output frame can be an inpainted image in which new pixels are filled in the area to be restored (hole area) of the input frame.

[0069] In one embodiment, the electronic device can generate training data for training the inpainting model (400). The electronic device can train the inpainting model (400) by generating training data including a correct image, a mask representing an occluded area, and a masked image. Once the inpainting model (400) is trained, the electronic device can use the inpainting model to restore an image or video. For example, the electronic device can acquire a monoscopic image or video, generate an image or video from a different viewpoint, and inpaint it to generate a stereoscopic image or video.

[0070] FIG. 4b is a diagram illustrating a detailed architecture of the inpainting model of the present disclosure.

[0071] Referring to Fig. 4b, the detailed configuration and input / output data of the inpainting model are illustrated. In describing the present disclosure, it will be assumed that the frame is the t-th frame of the video. In addition, unless otherwise stated, the feature (F) corresponding to the t-th frame t ) is an image feature (F t img ) and depth features (F t depth ) is included.

[0072] The image encoder (410) is a frame to be restored (X t img ) and mask (M t ) is input as an image feature (F t img ) can be printed.

[0073] The depth encoder (420) stores the frame to be restored (X t img ), mask (M t ) and depth frame (X t depth ) is input to obtain depth features (F t depth ) can be printed.

[0074] Image features (F t img ) and depth features (F t depth ) is concatenated to the t-th frame (X t img ) corresponding to the input feature (F t ) can be.

[0075] The alignment module (430) maintains temporal consistency in video inpainting by using the output features of the previous frame ( ) and the input features (F) of the current frame t ) can perform alignment. The alignment module (430) performs alignment of the aligned features (F) with new feature values. t TA ) can be printed.

[0076] The inpainting module (440) is used to sort the aligned features (F t TA ) containing the restored pixels of the current frame's output features. ) can be output. The inpainting module (440) outputs the aligned depth features (F t TA_depth ), the aligned image features (F t TA_img ) can be used to calculate the feature values. As a result, the output features ( ) may contain restored pixel information. Here, the aligned depth features (F t TA_depth ) may contain restored depth information. The aligned depth features (F t TA_depth ) is restored by the depth decoder (442) to create a depth feature (F t TA_depth ) is restored as a depth image, and the output depth frame ( ) can be evaluated by the inpainting model. The inpainting model outputs depth frames { } and the correct depth frame (Y t depth ) loss between (L) depth ), the depth encoder (420) generates depth features (F t depth ) can be included to include restored depth information when outputting.

[0077] The image decoder (450) outputs features ( ) to convert the output frame ( ) can be generated. Output frame ( ) is the input frame, which is the frame to be restored (X t img ) may be an inpainted image, with new pixels filled in the area to be restored (the hole area).

[0078] FIG. 5 is a diagram for explaining an operation of an electronic device according to one embodiment of the present disclosure to align a current frame with reference to a previous frame.

[0079] In one embodiment, the alignment module (500) of the inpainting model may include an attention module (510) and a Convolutional Gated Recurrent Unit (Conv GRU) (520).

[0080] In video inpainting, maintaining temporal consistency between the restored frame and the previous frame is crucial. To ensure temporal consistency, electronic devices can utilize information from previous frames when restoring the current frame. If the electronic device wishes to restore the current frame by referencing the previous frame, temporal alignment between the previous frame and the current frame is required. Furthermore, to ensure temporal alignment, spatial alignment of the previous frame is necessary to prevent spatial coordinate misalignment between frames. For this purpose, the attention module (510) and the Convolutional GRU (520) of the alignment module (510) can be used.

[0081] The attention module (510) can perform feature-to-feature alignment based on the attention mechanism. The attention module (510) uses the depth feature (Ftdepth) and the output feature of the previous frame ( ) as input, and the output features of the sorted previous frame ( ) can be output. The output features of the sorted previous frame ( ) is the output feature of the previous frame, frame t-1 ( ) can mean a feature that aligns spatial coordinates by spatially aligning the current frame, t frame.

[0082] The attention module (510) outputs the features of the previous frame ( ) are spatially aligned, and the output features of the (spatially) aligned previous frame ( ) can be obtained. As spatial alignment is performed by the attention module (510), when the Conv GRU (520) aligns the current frame by referring to the previous frame for temporal consistency, abnormal phenomena (e.g., frame afterimages, etc.) that may occur due to misalignment of spatial coordinates can be removed.

[0083] Conv GRU (520) outputs the sorted previous frame's output features ( ) and input features (F) of the current frame t ), based on the temporally aligned input features (F t TA ) can be obtained. The convolution layer of the Conv GRU (520) performs a convolution operation on the features provided as input, and the gated recurrent unit (GRU) can receive the output of the convolution operation as input and update the hidden state. Based on the hidden state, the GRU obtains the final output, the temporally aligned input feature (F t TA ) can be obtained. Temporally aligned input features (F t TA ) is a temporally aligned image feature (F t TA_img ) and temporally aligned depth features (F t TA_depth ) may be included.

[0084] When spatial alignment and temporal alignment of features are completed, the alignment module (500) can transfer the aligned features to the inpainting module.

[0085] Meanwhile, in explaining FIG. 5, for the sake of convenience, an example was given where the previous frame is the frame (t-1) immediately preceding the current frame (t). However, the embodiments of the present disclosure are not limited to the aforementioned example. For example, the previous frame may be a frame (tn) that is a predetermined number of frames prior to the current frame (t).

[0086] FIG. 6 is a diagram for explaining an operation of an electronic device according to one embodiment of the present disclosure to align a current frame with reference to a previous frame.

[0087] In one embodiment, the attention module of the inpainting model can perform feature-to-feature alignment based on an attention mechanism. The attention module can use the output features of the previous frame ( ) are spatially aligned, and the output features of the aligned previous frame ( ) can be obtained.

[0088] In the operation of the attention module, the depth feature (F) of the current frame t depth ) corresponds to the query, and the depth features of the previous frame ( ) corresponds to the key, and the image features of the previous frame ( ) corresponds to the value. In this case, the attention module can perform preprocessing to unfold the sequence-type features into a window shape.

[0089] The attention module can compute the similarity between the query and the key. The attention module uses the depth features (F) of the current frame t depth ) and the depth features of the previous frame ( ) can be calculated within the range r. Here, r is a parameter that sets the range for similarity calculation, meaning that the movement of the object and / or camera is within the range r in the feature space. Similarity calculation can be performed using, for example, the cosine similarity method, but is not limited thereto.

[0090] The attention module can calculate sorting weights, which are weights for sorting, by applying a softmax function to the calculated similarity. The weights calculated within the range r are: Wow F t depth The higher the similarity, the higher the weight value.

[0091] The attention module can apply weights to the values. The attention module can apply the alignment weights to the image features of the previous frame ( ), applying the image features of the previous frame ( ) aligned to the spatial coordinates of the current frame. ) can be obtained.

[0092] Meanwhile, the attention module, unlike the general attention mechanism, can apply weights not only to values ​​but also to keys. The attention module applies weights based on the depth features of the previous frame ( ), applying the depth feature of the previous frame ( ) is a depth feature aligned to the spatial coordinates of the current frame. ) can be obtained. The aligned depth features ( ) is the aligned image feature ( ) and aligned features ( ) can be configured. Sorted features ( ) can be used in the inpainting process of the inpainting module.

[0093] FIG. 7 is a diagram for explaining an operation of an electronic device inpainting an image according to one embodiment of the present disclosure.

[0094] In one embodiment, the inpainting module (700) of the inpainting model may include a depth-aware attention module (710) and a depth decoder (720).

[0095] The inpainting module (700) sorts the input features (F) from the sorting module. t TA ) can be passed. The sorted input features (F t TA ) is the aligned depth feature (F t TA_depth ) and aligned image features (F t TA_img ) can be separated into aligned depth features (F tTA_depth ) may contain restored depth information. The aligned depth features (F t TA_depth ) is restored by the depth decoder (720) into a depth feature (F t TA_depth ) is restored as a depth image, and the output depth frame ( ) can be verified by the inpainting model. The inpainting model outputs the depth frame ( ) and the correct depth frame (Y t depth ) loss between (L) depth ), the depth encoder generates depth features (F t depth ) can be included to include restored depth information when outputting.

[0096] The depth-aware attention module (710) uses the aligned depth features (F t TA_depth ) based on the restored depth information, the aligned image features (F t TA_img ) to calculate the feature values ​​of the output feature ( ) can be created.

[0097] The image decoder takes the sorted input features (F t TA ) of aligned image features (F t TA_img ) and output features ( ) image features( ), using at least one of the output frames ( ) can be generated. Output frame ( ) is the input frame, which is the frame to be restored (X t img ) may be an inpainted image, with new pixels filled in the area to be restored (the hole area).

[0098] FIG. 8 is a diagram illustrating a depth feature generated by an electronic device according to one embodiment of the present disclosure.

[0099] Referring to Fig. 8, the input image (810) is an image to be restored that includes a hole area. In addition, the input depth image (820) is an image representing depth information of the input image (810), and, like the input image (810), includes a hole area.

[0100] In one embodiment, when the inpainting model generates a depth feature using a depth encoder, the depth feature may include restored depth information. The restored depth information may be depth information filled in such that a hole area of ​​the input depth image (820) becomes the depth of a background area behind an object. When the electronic device trains the inpainting model, the depth encoder may train the inpainting model to generate a depth feature including the restored depth information. When the depth feature is restored using the depth decoder (800), an inpainted depth image (830) may be generated. The inpainted depth image (830) may refer to a depth image in which depth information is filled in a hole area.

[0101] In one embodiment, the inpainting model can first inpaint depth information and then inpaint an image based on the inpainted depth information. The operation of the inpainting model performing inpainting based on depth information is further described with reference to FIG. 9.

[0102] FIG. 9 is a diagram for explaining an operation of an electronic device restoring pixel information according to an embodiment of the present disclosure.

[0103] In one embodiment, the depth-aware attention module of the inpainting model can calculate weights for image inpainting based on an attention mechanism. The depth-aware attention module can restore pixel information based on restored depth information included in the depth feature.

[0104] In the operation of the depth-aware attention module, the depth feature (F) of the current frame t depth) corresponds to the key query. And, the image feature (F) of the current frame t img ) corresponds to the value. In this case, the depth-aware attention module can perform preprocessing to unfold the sequence-shaped features into a window shape for operation.

[0105] The depth-aware attention module can compute the similarity between a query and a key. The depth-aware attention module calculates the depth features (F) of the current frame. t depth ) can perform a similarity calculation by comparing which part and depth are similar within the range r. Here, r is a parameter that sets the range for similarity calculation, which means that it is assumed that the range in which occlusion caused by the object occurs is within the maximum range r. The similarity calculation can use, for example, the cosine similarity method, but is not limited thereto.

[0106] The depth-aware attention module can calculate inpainting weights, which are weights for inpainting, by applying a softmax function to the calculated similarity. The weights calculated based on the similarity within the r range have higher values ​​as the depth similarity increases. Accordingly, when pixel information is restored based on the weights, pixel information with similar depth information within the r range can be referenced.

[0107] The depth-aware attention module can apply weights to values. The depth-aware attention module applies the inpaint weights to the image features (F t img ), and the output image features ( containing the restored pixel information) are applied. ) can be obtained.

[0108] Meanwhile, the depth-aware attention module, unlike the general attention mechanism, can apply weights not only to values ​​but also to keys. The attention module applies weights to the depth features (F) of the current frame. t depth ), the output depth feature ( ) can be obtained. Output depth features ( ) is the output image feature ( ) and is attached to the output feature ( ) can be configured. Output features ( ) can be used when inpainting the next frame of the current frame. For example, the output feature of the current frame ( ) can be used to achieve spatial and temporal alignment with the next frame when inpainting the next frame.

[0109] FIG. 10 is a drawing for explaining an operation of an electronic device inpainting an image according to one embodiment of the present disclosure.

[0110] In one embodiment, the electronic device can inpaint an image. If the electronic device inpaints an image other than a video, the operation of the alignment module described above, which performs spatial and temporal alignment with previous frames, may be omitted.

[0111] The depth encoder (1010) stores the image to be restored (X t img ), mask (M t ) and depth image (X t depth ) is input to obtain depth features (F t depth ) can be printed.

[0112] The image encoder (1020) is an image to be restored (X t img ) and mask (M t ) is input as an image feature (F t img ) can be printed.

[0113] The depth-aware attention module (1030) can apply an attention mechanism based on depth information. The depth-aware attention module (1030) can apply a depth feature (F t depth ) and image features (F t img ) as input and output features ( ) can be output. The output feature is an output depth feature ( ) and output image features containing restored pixel information ( ) can be composed of.

[0114] The image decoder (1040) outputs features ( ) to convert the output frame ( ) can be generated. Output frame ( ) is the input frame, which is the frame to be restored (X t img ) may be an inpainted image, with new pixels filled in the area to be restored (the hole area).

[0115] In one embodiment, when an electronic device performs video inpainting, the first frame of the video has no previous frames. Therefore, the first frame of the video inpainting can be performed in the manner of image inpainting, omitting the operation of the alignment module, as illustrated in FIG. 10.

[0116] FIG. 11 is a diagram illustrating an operation of an electronic device generating training data according to an embodiment of the present disclosure.

[0117] In one embodiment, an electronic device can generate training data for training an inpainting model. In the conversion of images / videos from mono to stereo, the lack of stereo datasets with ground truth disparity presents challenges in training an inpainting model. To address the aforementioned data shortage, the electronic device can directly generate masked images and masks that include the restoration target region (the hole region) and use them as training data.

[0118] In one embodiment, the electronic device can generate an arbitrary disparity map based on depth information of an input image. The operation of the electronic device generating a disparity map based on depth information has been described above with reference to FIG. 3, and therefore, a repeated description thereof will be omitted.

[0119] The electronic device can warp and dewarp an input image based on disparity and inverse disparity.

[0120] For example, an electronic device can warp an input image to the left based on disparity, and then warp the input image back to the right (original position) based on inverse disparity. This can result in the creation of a hole region in the right area of ​​an object and the left area of ​​a background in the input image. The electronic device can then generate a mask (MaskLeft→Input) representing the right area of ​​the object and the left area of ​​the background.

[0121] Similarly, although not illustrated in FIG. 11, the electronic device can warp the input image to the right based on the inverse disparity and warp the input image to the left (original position) based on the disparity. Accordingly, a hole region can be created in the left area of ​​the object and the right area of ​​the background in the input image. The electronic device can generate a mask (MaskRight→Input) representing the left area of ​​the object and the right area of ​​the background.

[0122] The electronic device can generate a mask by combining the generated masks (MaskLeft→Input, MaskRight→Input) and apply the mask to the input image to generate a masked image. The masked image and mask can be used as training data for the inpainting model.

[0123] In one embodiment, an electronic device can train an inpainting model using a training dataset. The training dataset may include training data generated by the electronic device and other training data (e.g., human-generated data). The training dataset may include original images, masked images, masks, and depth images. In the aforementioned embodiments, the inpainting model used by the electronic device may refer to an inpainting model that has been deployed after training using the training dataset and performance verification and evaluation.

[0124] FIG. 12 is a drawing for explaining an embodiment of an electronic device performing an inpainting operation according to an embodiment of the present disclosure.

[0125] The above-described embodiments have been described based on examples in which an electronic device generates a new image or video from a new perspective through inpainting, or generates a stereo image / video, based on an image or video. However, the application of the inpainting method of the present disclosure is not limited to this.

[0126] In one embodiment, the electronic device may generate an image or video showing an object's position shifted in the image or video. For example, the electronic device may generate an image or video showing an object shifted from a first position (1210) to a second position (1220).

[0127] An electronic device may receive user input for selecting an object within an image. The user input may include various types of input for identifying an object, such as dragging, touching, swiping, or voice input. Alternatively, the electronic device may provide a feature that automatically detects separable objects within an image and allows the user to select them.

[0128] The electronic device may receive user input to edit a selected object within an image. User input may include, but is not limited to, drag input to move and / or resize an object, voice input indicating the direction of movement, distance moved, or resize of an object, text input, etc.

[0129] An electronic device can modify an image to generate a restored image, once an object is identified within the image and the final position of the selected object is identified. The restored image may contain an object moved or resized to a different position corresponding to user input, and an existing object may be repositioned into a hole area.

[0130] An electronic device can inpaint an image to be restored using an inpainting model. In this case, since the hole area is an occluded area previously covered by an object, the inpainting model can use depth information to reference pixels in the background area for inpainting. The inpainting model's operation on an image can omit the alignment module's operation and only include the inpainting module's operation. Since a detailed description of image inpainting has been provided above, a repetitive description will be omitted.

[0131] In one embodiment, the electronic device can generate frames to be reconstructed by modifying frames of the video. For example, the electronic device can receive user input for changing the size and / or position of an object in the first frame of the video, and generate an initial frame to be reconstructed that includes a hole area. Based on the initial frame to be reconstructed, the electronic device can modify subsequent frames to generate subsequent frames to be reconstructed.

[0132] An electronic device can inpaint frames to be restored using an inpainting model. The operation of the inpainting model for video may include the operation of an alignment module and an inpainting module. A detailed description of video inpainting has been provided above, so a repetitive description will be omitted.

[0133] FIG. 13 is a drawing for explaining an embodiment of an electronic device performing an inpainting operation according to an embodiment of the present disclosure.

[0134] In one embodiment, the electronic device can generate 3D motion effects based on images or videos.

[0135] For example, the electronic device can generate a 3D motion effect, such as a clockwise rotation of a viewpoint within the original image based on the original image (1300). Here, the clockwise rotation of the viewpoint within the image refers to a 3D effect in which the scene captured by the camera rotates clockwise as the camera moves counterclockwise from the point at which the original image (1300) was captured.

[0136] In one embodiment, an electronic device may receive user input to generate a 3D motion effect. The electronic device may receive user input to adjust perspective within a scene, adjust 3D rotation, etc. Alternatively, the electronic device may provide a function to automatically generate an arbitrary 3D motion effect based on an image.

[0137] The electronic device may generate a plurality of reconstructed images based on the original image (1300) to create a three-dimensional motion effect. For example, if the viewpoint within the original image (1300) rotates clockwise, the foreground (e.g., objects) within the image should move to the left, and the background (e.g., buildings) should move to the right. The electronic device may generate a plurality of reconstructed images to gradually display the three-dimensional rotation effect. For example, a first reconstructed image (1310), a second reconstructed image (1320), and a third reconstructed image (1330) may be generated. In this case, the degree of movement of the object and the background may increase from the first reconstructed image (1310) to the third reconstructed image (1330). Accordingly, the size of the hole area representing the occlusion area that occurs as the object and the background move in different directions may also increase from the first reconstructed image (1310) to the third reconstructed image (1330).

[0138] Electronic devices can inpaint images to be restored using an inpainting model. In this case, since the hole area is an occluded area previously covered by an object, the inpainting model can use depth information to reference pixels in the background area for inpainting. The operation of the inpainting model to generate a 3D motion effect may include both the alignment module and the inpainting module. Since a detailed description of image inpainting has been provided above, a repetitive description will be omitted.

[0139] FIG. 14 is a block diagram illustrating a configuration of an electronic device according to one embodiment of the present disclosure.

[0140] Referring to FIG. 14, an electronic device (2000) according to one embodiment may include a communication interface (2100), a memory (2200), a processor (2300), and a display (2400).

[0141] The communication interface (2100) can perform data communication with other electronic devices under the control of the processor (2300).

[0142] The communication interface (2100) may include a communication circuit that can perform data communication between the electronic device (2000) and another electronic device (e.g., a server (3000)) using at least one of data communication methods including, for example, wired LAN, wireless LAN, Wi-Fi, Bluetooth, ZigBee, Wi-Fi Direct (WFD), infrared Data Association (IrDA), Bluetooth Low Energy (BLE), Near Field Communication (NFC), Wireless Broadband Internet (Wibro), World Interoperability for Microwave Access (WiMAX), Shared Wireless Access Protocol (SWAP), Wireless Gigabit Alliances (WiGig), and RF communication.

[0143] The communication interface (2100) can transmit and receive data for performing image inpainting to and from an external device. For example, modules stored in the memory (2200) may be components of an inpainting model and downloaded from an external device (e.g., a server (3000)).

[0144] The memory (2200) may store instructions, data structures, and program codes that can be read by the processor (2300). Operations performed by the processor (2300) may be implemented by executing instructions or codes of a program stored in the memory (2200).

[0145] The memory (2200) may include a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD or XD memory, etc.), and may include a non-volatile memory including at least one of a ROM (Read-Only Memory), an EEPROM (Electrically Erasable Programmable Read-Only Memory), a PROM (Programmable Read-Only Memory), a magnetic memory, a magnetic disk, and an optical disk, and a volatile memory such as a RAM (Random Access Memory) or an SRAM (Static Random Access Memory).

[0146] The memory (2200) may store one or more instructions and / or programs that cause the electronic device (2000) to operate to provide image and / or video inpainting. For example, the memory (2200) may include an alignment module (2210), an inpainting module (2220), and a data augmentation module (2230). Meanwhile, the memory (2200) may further store instructions and / or programs for implementing functions of the image / video inpainting model. For example, the memory (2200) may store an image encoder / decoder, a depth encoder / decoder, and the like.

[0147] The processor (2300) can control the overall operations of the electronic device (2000). For example, the processor (2300) can control the overall operations of the electronic device (2000) for performing inpainting by executing one or more instructions of a program stored in the memory (2200). There may be one or more processors (2300).

[0148] The processor (2300) may be configured with at least one of, for example, a central processing unit, a microprocessor, a graphic processing unit, an application specific integrated circuits (ASICs), a digital signal processor (DSPs), a digital signal processing device (DSPDs), a programmable logic device (PLDs), a field programmable gate array (FPGAs), an application processor, a neural processing unit, or an artificial intelligence processor designed with a hardware structure specialized for processing an artificial intelligence model, but is not limited thereto.

[0149] The processor (2300) can execute the alignment module (2210) to perform frame-to-frame alignment operations required for video inpainting. Since the operations of the alignment module (2210) have already been described in the descriptions of the previous drawings, a repetitive description will be omitted.

[0150] The processor (2300) can execute an inpainting module (2220) to perform image inpainting or video frame inpainting. Since the operations of the inpainting module (2220) have already been described in the description of the previous drawings, a repetitive description will be omitted.

[0151] The processor (2300) can generate training data for training an inpainting model by executing a data augmentation module (2230). The operation of the data augmentation module (2230) has already been described in the descriptions of the previous drawings, and therefore, a repetitive description will be omitted.

[0152] Meanwhile, the modules stored in the aforementioned memory (2200) and executed by the processor (2300) are for convenience of explanation and are not necessarily limited thereto. To implement the aforementioned embodiments, other modules may be added, and some modules may be omitted. Furthermore, a single module may be divided into multiple modules distinguished by their detailed functions, and some of the aforementioned modules may be combined to be implemented as a single module.

[0153] When a method according to an embodiment of the present disclosure includes a plurality of operations, the plurality of operations may be performed by one processor (2300) or by a plurality of processors (2300). For example, when a first operation, a second operation, and a third operation are performed by a method according to an embodiment, the first operation, the second operation, and the third operation may all be performed by a first processor, or the first operation and the second operation may be performed by a first processor (e.g., a general-purpose processor) and the third operation may be performed by a second processor (e.g., an AI-dedicated processor). Here, an AI-dedicated processor, which is an example of the second processor, may perform operations for training / inference of an AI model. However, the embodiments of the present disclosure are not limited thereto.

[0154] One or more processors (2300) according to the present disclosure may be implemented as a single-core processor or as a multi-core processor.

[0155] When a method according to one embodiment of the present disclosure includes multiple operations, the multiple operations may be performed by one core or may be performed by multiple cores included in one or more processors (2300).

[0156] The display (2400) can output a video signal to the screen of the electronic device (2000) under the control of the processor (2300). For example, the display (2400) can output a video signal generated as the electronic device (2000) performs inpainting, such as an original image, an inpainted image, a result of a viewpoint conversion of an image, a result of an object movement within an image, a 3D motion effect, and a stereo video, on the screen. The display (2400) can include a touch panel. The touch panel can include one or more touch sensors that detect a touch input. In one embodiment, the prompt can be a text input input through the touch panel.

[0157] FIG. 15 is a block diagram illustrating the configuration of a server according to one embodiment of the present disclosure.

[0158] In one embodiment, the operations of the electronic device (2000) described above may be performed in a server (3000). The server (3000) may include a communication interface (3100), a memory (3200), and a processor (3300). The server (3000) may be a computing device with higher performance than the electronic device (2000), capable of processing complex operations and tasks using large amounts of data, such as training, inference, management, distribution, and operation of an inpainting model.

[0159] The communication interface (3100) may include a communication circuit that can perform data communication between the server (3000) and another electronic device (e.g., the electronic device (2000)) using at least one of data communication methods including, for example, wired LAN, wireless LAN, Wi-Fi, Bluetooth, ZigBee, Wi-Fi Direct (WFD), infrared Data Association (IrDA), Bluetooth Low Energy (BLE), Near Field Communication (NFC), Wireless Broadband Internet (Wibro), World Interoperability for Microwave Access (WiMAX), Shared Wireless Access Protocol (SWAP), Wireless Gigabit Alliances (WiGig), and RF communication.

[0160] The communication interface (3100) can transmit and receive data for inpainting to and from the electronic device (2000) under the control of the processor (3300). For example, the server (3000) can receive an inpainting request from the electronic device (2000) through the communication interface (3100) and transmit the inpainting result to the electronic device (2000).

[0161] The memory (3200) may store one or more instructions and / or programs that cause the server (3000) to operate to provide image and / or video inpainting. For example, the memory (3200) may include an alignment module (3210), an inpainting module (3220), and a data augmentation module (3230). Meanwhile, the memory (3200) may further store instructions and / or programs for implementing functions of the image / video inpainting model. For example, the memory (3200) may store an image encoder / decoder, a depth encoder / decoder, and the like.

[0162] The processor (3300) can control the overall operations of the server (3000). For example, the processor (3300) can control the overall operations of the server (3000) for performing inpainting by executing one or more instructions of a program stored in the memory (3200). There may be one or more processors (3300).

[0163] The processor (3300) may be configured as at least one of, for example, a Central Processing Unit, a microprocessor, a Graphic Processing Unit, Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), an Application Processor, a Neural Processing Unit, or an artificial intelligence processor designed with a hardware structure specialized for processing artificial intelligence models, but is not limited thereto.

[0164] The processor (3300) can perform inpainting tasks by executing modules stored in memory. This has already been described in the description of the previous drawings, so a repeated description will be omitted.

[0165] The present disclosure relates to a method, electronic device, and server for inpainting images and / or videos. The present disclosure relates to an inpainting method for restoring an area occluded by an object based on depth information. The technical challenges addressed by the present disclosure are not limited to those described above, and other technical challenges not mentioned herein will be readily apparent to those skilled in the art, based on the description herein.

[0166] According to one aspect of the present disclosure, a method for an electronic device to inpaint a video may be provided.

[0167] The method may include a step of extracting input features including image features and depth features based on a first frame to be restored, a mask, and a depth frame.

[0168] The method may include selecting a second frame that is a frame prior to the first frame and is associated with the first frame.

[0169] The method may include a step of aligning an input feature of the first frame with an output feature of the second frame based on depth features of the first frame and the second frame.

[0170] The above method may include a step of obtaining an output feature by modifying an image feature based on a depth feature of the sorted input feature.

[0171] The method may include a step of generating a restored frame based on the output feature.

[0172] The above method may include a step of estimating the depth of the original frame.

[0173] The method may include a step of obtaining the first frame to be restored by converting the viewpoint of the original frame based on depth information of the original frame.

[0174] The step of extracting the above input feature may include a step of extracting an image feature using the first frame and a mask.

[0175] The step of extracting the input feature may include a step of extracting a depth feature using the first frame, the mask, and the depth frame.

[0176] The above depth feature may include depth information that restores a hole area corresponding to the mask within the depth frame.

[0177] The step of aligning the input feature of the first frame with the output feature of the second frame may include the step of spatially aligning the output feature of the second frame with the first frame based on the depth feature of the second frame and the depth feature of the first frame.

[0178] The step of aligning the input features of the first frame with the output features of the second frame may include the step of temporally aligning the aligned output features of the second frame and the input features of the first frame.

[0179] The step of spatially aligning the output feature of the second frame with the first frame may include the step of calculating a similarity between a specific pixel of the depth feature of the first frame and surrounding pixels of the depth feature of the second frame.

[0180] The step of spatially aligning the output features of the second frame to the first frame may include a step of calculating alignment weights based on the similarity.

[0181] The step of spatially aligning the output features of the second frame with the first frame may include the step of applying the alignment weights to the image features of the second frame.

[0182] The step of obtaining the above output feature may include a step of calculating a similarity between a specific pixel and surrounding pixels with respect to a depth feature of the sorted input feature.

[0183] The step of obtaining the above output feature may include a step of calculating an inpaint weight based on the above similarity.

[0184] The step of obtaining the above output feature may include the step of applying the inpaint weight to the image feature of the first frame.

[0185] The method may include a step of generating training data for video inpainting.

[0186] The step of generating the training data may include the step of warping and dewarping an image to generate a mask representing an area occluded by an object in the image.

[0187] The step of generating the training data may include a step of generating the training data comprising the correct image, a mask representing the masked area, and a masked image.

[0188] The method may include a step of generating a stereo video including the original frame and the restored frame.

[0189] The method may include a step of generating a plurality of restored frames to generate a video including a three-dimensional motion effect.

[0190] According to one aspect of the present disclosure, an electronic device for performing inpainting may be provided. The electronic device may include a memory storing one or more instructions; and one or more processors for executing the one or more instructions stored in the memory.

[0191] The one or more processors can extract input features including image features and depth features based on the first frame to be restored, the mask, and the depth frame by executing the one or more instructions.

[0192] The one or more processors can select a second frame that is a frame prior to the first frame and associated with the first frame by executing the one or more instructions.

[0193] The one or more processors can align an input feature of the first frame with an output feature of the second frame based on depth features of the first frame and the second frame by executing the one or more instructions.

[0194] The one or more processors can obtain output features by modifying image features based on depth features of the sorted input features by executing the one or more instructions.

[0195] The one or more processors can generate a restored frame based on the output feature by executing the one or more instructions.

[0196] The one or more processors can estimate the depth of the original frame by executing the one or more instructions.

[0197] The one or more processors can obtain the first frame to be restored by converting the viewpoint of the original frame based on the depth information of the original frame by executing the one or more instructions.

[0198] The one or more processors can extract image features using the first frame and the mask by executing the one or more instructions.

[0199] The one or more processors can extract depth features using the first frame, the mask, and the depth frame by executing the one or more instructions.

[0200] The above depth feature may include depth information that restores a hole area corresponding to the mask within the depth frame.

[0201] The one or more processors can spatially align an output feature of the second frame with respect to the first frame based on the depth feature of the second frame and the depth feature of the first frame by executing the one or more instructions.

[0202] The one or more processors can temporally align the aligned output features of the second frame and the input features of the first frame by executing the one or more instructions.

[0203] The one or more processors can calculate a similarity between a specific pixel of the depth feature of the first frame and surrounding pixels of the depth feature of the second frame by executing the one or more instructions.

[0204] The one or more processors can calculate alignment weights based on the similarity by executing the one or more instructions.

[0205] The one or more processors can apply the alignment weights to the image features of the second frame by executing the one or more instructions.

[0206] The one or more processors can calculate a similarity between a specific pixel and surrounding pixels with respect to a depth feature of the sorted input feature by executing the one or more instructions.

[0207] The one or more processors can calculate inpaint weights based on the similarity by executing the one or more instructions.

[0208] The one or more processors can apply the inpaint weights to the image features of the first frame by executing the one or more instructions.

[0209] The one or more processors can generate training data for video inpainting by executing the one or more instructions.

[0210] The one or more processors can warp and dewarp an image by executing the one or more instructions, thereby generating a mask representing an area occluded by an object in the image.

[0211] The one or more processors can generate the training data comprising the correct image, a mask representing the masked area, and a masked image by executing the one or more instructions.

[0212] The one or more processors can generate a stereo video including the original frame and the restored frame by executing the one or more instructions.

[0213] The one or more processors can generate a stereo video including the original frame and the restored frame by executing the one or more instructions, and generate a plurality of the restored frames to generate a video including a three-dimensional motion effect.

[0214] Meanwhile, embodiments of the present disclosure may also be implemented in the form of a recording medium containing computer-executable instructions, such as program modules, executed by a computer. Computer-readable media may be any available media that can be accessed by a computer, and include both volatile and nonvolatile media, removable and non-removable media. Furthermore, computer-readable media may include computer storage media and communication media. Computer storage media include both volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Communication media may typically include computer-readable instructions, data structures, or other data in a modulated data signal, such as program modules.

[0215] Additionally, a computer-readable storage medium may be provided in the form of a non-transitory storage medium. Here, the term "non-transitory storage medium" simply means a tangible device that does not contain signals (e.g., electromagnetic waves). This term does not distinguish between cases where data is permanently stored in the storage medium and cases where data is temporarily stored. For example, a "non-transitory storage medium" may include a buffer in which data is temporarily stored.

[0216] According to one embodiment, the method according to various embodiments disclosed in the present document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., a downloadable app) may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

[0217] The above description of the present disclosure is provided for illustrative purposes only, and those skilled in the art will readily appreciate that modifications to other specific forms can be made without altering the technical spirit or essential features of the present disclosure. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive. For example, components described as being single may be implemented in a distributed manner, and similarly, components described as being distributed may be implemented in a combined manner.

[0218] The scope of the present disclosure is indicated by the claims described below rather than the detailed description above, and all changes or modifications derived from the meaning and scope of the claims and their equivalent concepts should be interpreted as being included in the scope of the present disclosure.

Claims

1. In the video inpainting method, A step of extracting input features including image features and depth features based on a first frame to be restored, a mask, and a depth frame; A step of selecting a second frame, said second frame being a frame prior to said first frame and associated with said first frame; A step of aligning an input feature of the first frame with an output feature of the second frame based on the depth features of the first frame and the second frame; A step of obtaining an output feature by modifying an image feature based on the depth feature of the above-mentioned sorted input feature; and A method comprising the step of generating a restored frame based on the above output features.

2. In paragraph 1, The above method, A step of estimating the depth of the original frame; and A method further comprising the step of obtaining the first frame to be restored by transforming the viewpoint of the original frame based on depth information of the original frame.

3. In paragraph 1, The step of extracting the above input features is: A step of extracting image features using the first frame and mask; and A method comprising the step of extracting depth features using the first frame, the mask and the depth frame.

4. In paragraph 3, A method wherein the depth feature includes depth information that restores a hole area corresponding to the mask within the depth frame.

5. In paragraph 4, The step of aligning the input feature of the first frame with the output feature of the second frame is: A step of spatially aligning the output feature of the second frame to the first frame based on the depth feature of the second frame and the depth feature of the first frame; and A method comprising the step of temporally aligning the aligned output features of the second frame and the input features of the first frame.

6. In paragraph 5, The step of spatially aligning the output features of the second frame to the first frame is: A step of calculating the similarity between a specific pixel of the depth feature of the first frame and surrounding pixels of the depth feature of the second frame; A step of calculating alignment weights based on the above similarity; and A method comprising the step of applying the alignment weights to the image features of the second frame.

7. In paragraph 4, The step of obtaining the above output feature is: A step of calculating the similarity between a specific pixel and surrounding pixels for the depth feature of the above-mentioned sorted input features; A step of calculating inpaint weights based on the above similarity; and A method comprising the step of applying the inpaint weights to image features of the first frame.

8. In electronic devices, Memory that stores one or more instructions; and comprising one or more processors executing one or more instructions stored in said memory; The one or more processors, by executing the one or more instructions, Extract input features including image features and depth features based on the first frame to be restored, the mask, and the depth frame, Select a second frame that is a frame prior to the first frame and associated with the first frame, Based on the depth features of the first frame and the second frame, the input features of the first frame are aligned with the output features of the second frame, Obtain output features by modifying image features based on the depth features of the above-mentioned sorted input features, An electronic device for generating a restored frame based on the above output features.

9. In paragraph 8, The one or more processors, by executing the one or more instructions, Estimate the depth of the original frame, An electronic device that obtains the first frame to be restored by converting the viewpoint of the original frame based on depth information of the original frame.

10. In paragraph 8, The one or more processors, by executing the one or more instructions, Extract image features using the first frame and mask above, An electronic device for extracting depth features using the first frame, the mask and the depth frame.

11. In paragraph 10, An electronic device, wherein the depth feature includes depth information that restores a hole area corresponding to the mask within the depth frame.

12. In paragraph 11, The one or more processors, by executing the one or more instructions, Spatially aligning the output feature of the second frame to the first frame based on the depth feature of the second frame and the depth feature of the first frame; An electronic device that temporally aligns the aligned output features of the second frame and the input features of the first frame.

13. In paragraph 12, The one or more processors, by executing the one or more instructions, Compute the similarity between a specific pixel of the depth feature of the first frame and surrounding pixels of the depth feature of the second frame, Calculate the alignment weights based on the above similarity, An electronic device that applies the alignment weights to the image features of the second frame.

14. In paragraph 11, The one or more processors, by executing the one or more instructions, For the depth feature of the above sorted input features, the similarity between a specific pixel and surrounding pixels is calculated, Calculate the inpaint weights based on the above similarity, An electronic device that applies the inpaint weights to image features of the first frame.

15. A computer-readable recording medium having recorded thereon a program for executing the method of any one of clauses 1 to 7 on a computer.

Citation Information

Patent Citations

  • Bolometer device and method of manufacturing the bolometer device

    KR102284749B1

  • Deep patch feature prediction for image inpainting

    US20190295227A1