Method and device for generating outer portion of video in order to change video aspect ratio
The method employs a mask-filling deep network and a repetitive outer extension rapper to ensure natural video aspect ratio conversion from 4:3 to 16:9, maintaining frame consistency and preventing object cutoffs.
Patent Information
- Application Number
- PCT/KR2024/016367
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-30
- Filing Date
- 2024-10-25
- Publication Date
- 2025-05-08
AI Technical Summary
Existing methods for changing the screen ratio of videos, such as from 4:3 to 16:9, often result in unnatural expansions of objects and can lead to objects being cut off or not generated properly, especially when maintaining consistency between frames.
A method and device using a mask-filling deep network to generate a restoration image by receiving an image and mask, along with a conventional predictor to predict the restoration image's convention, and a repetitive outer extension rapper to control the outer area of the video, ensuring objects are not cut off and the aspect ratio conversion is natural.
This approach maintains consistency between frames during video aspect ratio conversion, ensuring a natural and seamless transformation of the video's screen ratio without cutting off objects.
Smart Images

Figure KR2024016367_08052025_PF_FP_ABST
Abstract
Description
Method and device for generating the outer part of a video for changing the aspect ratio of the video
[0001] The present invention relates to a method and device for generating an outer portion of a video for changing the aspect ratio of the video, and more particularly, to a device and method for filling an image in an empty space in the outer portion of a video when changing the aspect ratio of the video.
[0002] In the past, the standard aspect ratio for TV broadcasting equipment was 4:3. Consequently, much video content was produced in this ratio. Recently, with the diversification of broadcasting platforms and the advancement of broadcasting equipment and display devices, videos are being produced in various aspect ratios. The most commonly used aspect ratio is 16:9.
[0003] Meanwhile, there's a growing need to replay or utilize past content. However, to play back content produced in the 4:3 ratio on today's widely used 16:9 displays, the video aspect ratio needs to be adjusted to 16:9.
[0004] One way to change the aspect ratio is to directly adjust the aspect ratio. For example, by stretching a 4:3 image horizontally to create a 16:9 aspect ratio. While this method makes it easy to change the aspect ratio, it causes objects within the image to expand horizontally, creating an unnatural effect.
[0005] As a method to solve this problem, a method of changing the ratio of both sides of the main object while maintaining the ratio of the main object is proposed.
[0006] Alternatively, a segmentation network can be used to separate the central object from the background in the input image, an inpainting network can be used to fill in the missing portion of the image, and then the background can be rescaled and the object reattached. This method maintains the proportions of the main object, resulting in a natural look.
[0007] Another approach, called outpainting, involves creating images of the outer edges (hereinafter referred to as "outer areas") of an existing 4:3 image without altering it, and then filling them in to create a 16:9 image. Recent advances in stable diffusion technology have led to outstanding performance of outpainting models.
[0008] However, changing the aspect ratio of a video is not simple. Unlike still images, video requires consistency in the areas generated between adjacent frames. For example, if there are differences between adjacent frames in areas such as a stretched background or an outpainted outer area, it can be visually jarring.
[0009] Furthermore, conventional outpainting methods, which separate objects and backgrounds, can sometimes generate only the background and not the object itself when the object is located on either side of the image and a portion of the object is cropped. This can lead to additional issues, such as the object appearing cropped in the outpainted image.
[0010] The present invention has been made in consideration of these points, and its purpose is to provide a method and device for generating an outer portion of a video for changing the aspect ratio of the video, which can maintain consistency between frames while converting the aspect ratio of the video.
[0011] Another object of the present invention is to provide a method and device for generating an outer portion of a video for changing the screen ratio of a video without cutting off objects located in the outer portion.
[0012] The problems to be solved by the present invention are not limited to the problems mentioned above, and other problems not mentioned will be clearly understood by those skilled in the art from the description below.
[0013] A device for generating an outer region of a video for changing an aspect ratio of a video according to a preferred embodiment of the present invention comprises: a mask filling deep network for receiving an image and a mask as input and generating a restored image; a confidence predictor for predicting confidence of the restored image; and an iterative outer region expansion wrapper for controlling filling of an outer region of the video through N iterative operations in the mask filling deep network and the confidence predictor when an input video of a first aspect ratio is given until an image of a second aspect ratio is generated.
[0014] In one embodiment, the iterative edge expansion wrapper inputs an image[i] and a mask[i] to the mask-filling deep network and runs the deep network. The mask-filling deep network outputs a reconstructed image[i] with the edge filled using the input image[i] and mask[i]. The confidence predictor outputs a confidence map[i] indicating the confidence of the filled region.
[0015] Mask[i] can be generated by performing a pixel-wise sum of maski and the confidence map[i-1] generated in the previous operation and applying thresholding. The reconstructed image[i-1] generated in the previous operation is used as image[i]. The iterative outer expansion wrapper repeats the above process from i=1 to i=N and then outputs the reconstructed image output from the mask filling deep network as the final output image.
[0016] In one embodiment, the iterative outer expansion wrapper controls the performing of N iterative operations while incrementally increasing the size of mask i between a first aspect ratio and a second aspect ratio.
[0017] In one embodiment, the output of the confidence predictor is a confidence map of the same size as the image size of the second aspect ratio. Each pixel in the confidence map may have a value of 0 or 1 based on whether the confidence value is greater than or equal to a predetermined value.
[0018] In one embodiment, the mask filling deep network comprises a flow completion unit that extracts and completes optical flow between adjacent frames from input images, a CNN encoder that encodes features of the input image, a feature propagation unit that calculates content to be filled in an outer portion of the image using the optical flow features generated by the flow completion unit and the context features encoded by the CNN encoder, a temporal transformer that performs an outer region generation task by combining information from adjacent and non-adjacent neighbors, and a CNN decoder that outputs a reconstructed image with an outer portion filled in from information synthesized by the temporal transformer. The output of the temporal transformer is input to the confidence predictor.
[0019] In one embodiment, the device for generating an outline of a video according to the present invention may be trained using a first mask used to train the device to have a function of generating an object by learning the optical flow of the object, and a second mask used to train the function of filling in the outline area. The first mask is a mask that segments an object located in the image, and the second mask is an outpainting mask.
[0020] A method for generating an outer portion of a video for changing an aspect ratio of the video according to a preferred embodiment of the present invention comprises a first step in which an iterative outer portion expansion wrapper inputs an input video having a first aspect ratio and an initial mask, mask 1, into a mask filling deep network, a second step in which the mask filling deep network outputs a reconstructed image whose outer portion is filled using the input image and the mask, and a confidence map indicating the confidence of the filled area is output by a confidence predictor, a third step in which a pixel-wise sum of mask i and a confidence map [i-1] generated in a previous operation is performed and thresholding is applied to generate mask [i], and the mask [i] and the reconstructed image [i-1] generated in the previous operation are input into the mask filling deep network, and a fourth step in which the iterative outer portion expansion wrapper is controlled to repeat the second and third steps a predetermined number of times to output a reconstructed image having a second aspect ratio size output from the mask filling deep network as a final output image.
[0021] In one embodiment, the iterative outer expansion wrapper can be controlled to perform N iterative operations while gradually increasing the size of mask i between a first aspect ratio and a second aspect ratio.
[0022] In one embodiment, the output of the confidence predictor may be a confidence map having a size equal to the image size of the second aspect ratio. The value of each pixel in the confidence map may have a value of 0 or 1 based on whether the confidence value is greater than or equal to a predetermined value.
[0023] A mask filling deep network comprises a flow completion unit that extracts and completes optical flow between adjacent frames from input images, a CNN encoder that encodes features of the input image, a feature propagation unit that calculates content to be filled in the outer portion of the image using the optical flow features generated by the flow completion unit and the context features encoded by the CNN encoder, a temporal transformer that performs an outer region generation task by combining information from adjacent and non-adjacent neighbors, and a CNN decoder that outputs a reconstructed image with the outer portion filled in from information synthesized by the temporal transformer. The output of the temporal transformer is input to the confidence predictor.
[0024] In one embodiment, the method for generating an outline of a video may further comprise a step of learning using a first mask used to learn the function of generating an object by learning the optical flow of the object, and a second mask used to learn the function of filling in the outline area. The first mask may be a mask that segments an object located in the image, and the second mask may be an outpainting mask.
[0025] According to the present invention, when converting the aspect ratio of a video, consistency between frames can be maintained, thereby enabling natural aspect ratio conversion.
[0026] According to the present invention, objects located on the periphery can be prevented from being cut off, thereby providing a more natural transformed image.
[0027] The effects of the present invention are not limited to the effects mentioned above, and other effects not mentioned will be clearly understood by those skilled in the art from the description below.
[0028] FIG. 1 is a block diagram showing the configuration of a device for generating an outer portion of a video according to one embodiment of the present invention.
[0029] FIG. 2 is a flowchart showing the operation flow of a method for generating an outer portion of a video according to one embodiment of the present invention.
[0030] FIG. 3 is a diagram showing a mask and a confidence map step by step according to one embodiment of the present invention.
[0031] FIG. 4 is an explanatory diagram for explaining the operation of a method for generating an outer portion of a video according to one embodiment of the present invention.
[0032] FIG. 5 is a block diagram showing the internal configuration of a mask filling deep network according to one embodiment of the present invention.
[0033] Figure 6 is a block diagram showing an example of a context feature encoder CNN network.
[0034] Figure 7 is a block diagram showing an example of SpyNet computing optical flow.
[0035] Figure 8 is a block diagram showing one example of a characteristic propagation unit.
[0036] Figure 9 is a block diagram showing one example of a temporal transformer module.
[0037] Figure 10 is a block diagram showing an example of a CNN decoder.
[0038] Figure 11 is a block diagram showing an example of a CNN confidence predictor.
[0039] Figure 12 is a diagram showing an example of an area requiring additional restoration predicted by the confidence map.
[0040] Figure 13 shows examples of masks used for learning, (a) showing an example of an inpainting mask and (b) showing an example of an outpainting mask.
[0041] Figure 14 shows the results of a quantitative evaluation experiment of the present invention.
[0042] Figure 15 shows another result of a quantitative evaluation experiment of the present invention.
[0043] Figure 16 shows a captured screen of an original video with a 4:3 ratio and a video expanded to a 16:9 ratio by applying the method of the present invention.
[0044] Figure 17 is a drawing showing the result when the method of the present invention is applied to a case where an object exists on the outskirts and the ratio is expanded to 16:9.
[0045] Figure 18 shows the result screen when the conventional method is applied and when the method of the present invention is applied.
[0046] Figure 19 shows a result screen when the method of the present invention is applied to animation.
[0047] Hereinafter, embodiments of the present invention will be described in detail with reference to exemplary drawings. When designating components in each drawing, it should be noted that, where possible, identical components are given the same reference numerals, even if they appear in different drawings. Furthermore, in describing the present embodiments, detailed descriptions of related known structures or functions will be omitted if they are deemed to obscure the gist of the present embodiments.
[0048] The following example illustrates converting a 4:3 ratio video to a 16:9 ratio, but the present invention is also applicable to conversion between other ratios.
[0049] When converting from a 4:3 ratio to a 16:9 ratio directly, the amount of content that must be generated for edge area expansion makes it difficult to efficiently generate results in one go. In this invention, the mask is gradually enlarged and some areas of the background are gradually expanded, starting from a small area. In other words, this invention proposes a video outpainting method that iteratively expands both horizontal ends of a 4:3 ratio video based on a confidence map while maintaining the original content.
[0050] FIG. 1 is a block diagram showing the configuration of a video outer portion generation device for changing the screen ratio of a video according to one embodiment of the present invention.
[0051] A video edge generation device (100) for changing the screen ratio of a video according to one embodiment of the present invention comprises an iterative edge extension wrapper (110), a mask fill deep network (120) for generating a restored image from an input image under the control of the iterative edge extension wrapper (110), and a confidence predictor (130) for predicting the confidence of the restored image. The confidence predictor (130) is used to indicate the restored quality of each pixel of the restored image.
[0052] The iterative outer expansion wrapper (110) controls the outer region of a video to be filled through N iterative operations until a 16:9 video is generated when an input video with a 4:3 ratio is given. This operation is described with reference to FIGS. 2 and 3. FIG. 2 is a flowchart showing the operation flow of a method for generating an outer region of a video according to an embodiment of the present invention, and FIG. 3 is a diagram showing a mask and a confidence map step by step according to an embodiment of the present invention.
[0053] In the present invention, N iterative operations are performed while gradually increasing the size of the mask from a 4:3 ratio to a 16:9 ratio. In Fig. 3 (a) to (c), mask 1 in step 1, mask 2 in step 2, and mask N in step N are illustrated, respectively. The example of Fig. 3 shows a case where a mask covering 3 / 4 of the screen is used in step 1, a mask covering 5 / 6 of the screen is used in step 2, and a mask covering 15 / 16 of the screen is used in step N.
[0054] In Fig. 2, step S10 is executed only in the first step (step 1) among N repeated operation steps, and in the subsequent operation steps, steps S20 to S50 are repeatedly executed.
[0055] When the operation is initiated, the iterative outer expansion wrapper (110) sets the initial data (step S10). That is, i is set to 1, the input video I with a 4:3 ratio is set to video [i], and the first mask, Mask 1 (see (a) of FIG. 3), is set to Mask [i]. In the following description, in order to distinguish between Mask 1 to Mask N shown in (a) to (c) of FIG. 3 and the mask input to the mask filling deep network (120), the mask input to the mask filling deep network (120) in the ith step is represented as Mask [i].
[0056] The iterative outer edge expansion wrapper (110) inputs the image [i] and the mask [i] into the mask filling deep network (120) and executes the deep network (step S20). The mask filling deep network (120) outputs a reconstructed image [1] whose outer edge is filled using the input image and mask, and the confidence predictor (130) outputs a confidence map [1] (see (f) of FIG. 3) indicating the confidence of the filled area. For example, when the size of an image with a 16:9 aspect ratio is HxW, the output of the confidence predictor (130) may be in the form of a map with a size of HxWx1. In one embodiment, the value of each pixel of the confidence map may have a value of 0 or 1 based on whether the confidence value is greater than or equal to a predetermined value. For example, 0 can be configured to represent a pixel whose confidence value is greater than or equal to a predetermined value, and 1 can be configured to represent a pixel whose confidence value is less than or equal to a predetermined value.
[0057] When the operation is finished, it is checked in step S30 whether the repeated operation has been executed N times. If the operation has been executed less than N times, the number of times count i is increased (step S40), and a pixel-wise sum of the mask i and the confidence map [i-1] generated in the previous operation is performed, and thresholding is applied to leave only those values greater than a predetermined threshold value, thereby generating a mask [i] used in the next step. Then, the restored image [i-1] generated in the previous operation is set as the image [i] input to the next step (step S40) and input to the mask filling deep network (120).
[0058] To illustrate this by taking the case of moving from step 1 to step 2 as an example, after the first operation is executed, i=2, and in step 2, pixel-wise summation and thresholding are performed on mask 2 (see (b) of FIG. 3) and the confidence map [1] generated in step 1 ((f) of FIG. 3). An example of the mask thus generated is illustrated in (d) of FIG. 3. This mask is input to the mask filling deep network (120) together with the restored image [1] generated in step 1 as mask [2]. The confidence map [2] generated in step 2 ((g) of FIG. 3) undergoes pixel-wise summation and thresholding with mask 3 to become the mask of step 3, and the restored image generated in step 2 becomes the input image of step 3. In the final step, step N, pixel-wise summation and thresholding are performed on the mask N ((c) of FIG. 3) and the confidence map [N-1] generated in step N-1 to generate the mask [N] ((e) of FIG. 3), and the mask [N] and the restored image [N-1] are input to the mask filling deep network (120).
[0059] In this way, steps S20 to S50 are repeatedly executed, and if i=N in step S30, step S60 is performed and the generated restored image (N) is output as the final result image.
[0060] The above operation is explained separately for each operation step with reference to FIG. 4. FIG. 4 is an explanatory diagram for explaining the operation of a method for generating an outer portion of a video according to one embodiment of the present invention.
[0061] In step 1, the iterative outer edge expansion wrapper (110) inputs the input image (I) and mask 1 (m1) to the mask filling deep network (120). The mask filling deep network (120) outputs a reconstructed image [1] (r1) with the outer edge filled using the input image and mask, and a confidence map [1] (c1) indicating the confidence of the filled area.
[0062] In step 2, pixel-wise summation and thresholding (111) are performed on the mask 2 (m2) and the confidence map [1] (c1) generated in step 1. The mask generated in this way is input to the mask filling deep network (120) as a mask [2] together with the restored image [1] (r1) generated in step 1. The mask filling deep network (120) outputs a restored image [2] (r2) with the outline filled using the input image and mask and a confidence map [2] (c2). The confidence map [2] (c2) and the restored image [2] (r2) generated in step 2 are used in step 3.
[0063] By repeating this operation, when we reach the final step, step N, the mask N(m N ) and the confidence map [N-1] generated in step N-1 are subjected to pixel-wise summation and thresholding (111) to generate a mask [N], and the mask [N] and the restored image [N-1] are input to the mask filling deep network (120). The mask filling deep network (120) uses the input image and mask to generate a restored image [N] (r) with the outer edges filled. N ) is output, and the restored image [N] is output as the final result image.
[0064] Next, the configuration of a mask fill deep network will be described with reference to FIGS. 5 to 10.
[0065] Conventional outpainting or retargeting methods generate only the background, resulting in unnatural clipping of objects at either end. This can be addressed by replacing them with a video inpainting model that generates both the background and the object. Video inpainting models are trained to erase objects within a video and restore the background.
[0066] For example, the mask-filling deep network (120) proposed by Zhen Li et al. 2 Although FGVI (End-to-End framework for Flow-Guided Video Inpainting) can be used (see Zhen Li, Cheng-Ze Lu, Jianhua Qin, Chun-Le Guo, Ming-Ming Cheng. Towards An End-to-End Framework for Flow-Guided Video Inpainting, InCVPR2022), the present invention is not limited to a specific mask-filling deep network.
[0067] E 2 FGVI learns the changes of objects based on optical flow, removes objects from images, and restores the background with high performance. In addition, because it has learned the flow of objects, it can also use the learned optical flow for a new purpose called object restoration.
[0068] FIG. 5 is a block diagram showing the internal configuration of a mask filling deep network (120) according to one embodiment of the present invention.
[0069] The flow completion unit (121) extracts and completes optical flow between adjacent frames from input images. The CNN (Convolutional Neural Network) encoder (122) encodes features of the input image (image [i]). The optical flow features generated by the flow completion unit and the context features encoded by the CNN encoder are input to the feature propagation module (123). The feature propagation unit (123) calculates content to be filled in the outer part of the image using the optical flow and the context features.
[0070] The temporal transformer (124) performs the task of generating an outer region by combining information from adjacent and non-adjacent neighbors. Filling in the outer region solely with information provided from adjacent frames is not sufficient. Content not shown in adjacent frames may appear in non-adjacent frames. Therefore, information from non-adjacent neighbors can be used as additional information for the missing regions in adjacent frames. Therefore, multiple temporal transformer blocks are applied to perform the task of generating an outer region by combining information from adjacent and non-adjacent neighbors.
[0071] The information synthesized by the temporal transformer (124) passes through the CNN decoder (125) and is restored into an image with the outer edges filled in (restored image [i]). In addition, the output of the temporal transformer (124) is input to the confidence predictor (130).
[0072] Below, each component of the mask filling deep network (120) is described in detail.
[0073] The CNN encoder (122) can be implemented using a context feature encoder CNN network. Fig. 6 is a block diagram showing an example of a context feature encoder CNN network. In one embodiment, the CNN encoder (122) encodes input features to reduce the amount of computation by changing an input frame of size HxW to size 1 / 4*H x 1 / 4*W.
[0074] The flow completion unit (121) calculates the optical flow between two adjacent frames (forward t->t+1, backward t->t-1). For example, SpyNet, as illustrated in FIG. 7, can be used as the flow completion unit (121), and the present invention is not limited to a specific flow completion unit. (See Anurag Ranjan and Michael J Black. Optical flow estimation using a spatial pyramid network. In CVPR, 2017.)
[0075] The feature propagation unit (123) calculates the content to be filled in the outer part of the image using optical flow and context characteristics. The optical flow provides information to ensure that adjacent frames are well aligned with each other. The feature propagation unit (123) obtains the information necessary to fill in the information of the current frame from the aligned adjacent frames. As the feature propagation unit (123), for example, the E proposed by Zhen Li et al., as shown in FIG. 8, 2 FGVI (End-to-End framework for Flow-Guided Video Inpainting) can be used, and the present invention is not limited to a specific characteristic propagation section.
[0076] Filling in the perimeter with only information provided from adjacent frames is not sufficient. Content not shown in adjacent frames may appear in non-adjacent frames. Therefore, information from non-adjacent neighbors can be used as additional information for the missing regions in adjacent frames. Therefore, multiple temporal transformers (124) are applied to combine information from adjacent and non-adjacent neighbors to generate the perimeter region. Examples of temporal transformers include the E proposed by Zhen Li et al., as shown in Fig. 9. 2 The temporal transformer presented in FGVI (End-to-End framework for Flow-Guided Video Inpainting) can be used, and the present invention is not limited to the temporal transformer.
[0077] The information synthesized in the temporal transformer (124) is restored as an image with the outer edges filled in by passing through the CNN decoder (125). Fig. 10 is a block diagram showing an example of a CNN decoder.
[0078] Next, referring to Fig. 11, a confidence predictor (130) is described. The confidence predictor (130) is used to indicate the restored quality of each pixel of the restored image. The output of the confidence predictor (130) is a map of size HxWx1, where the value c of each pixel is c∈[0,1], where 0 indicates a pixel restored with very good quality and 1 indicates a pixel with poor restored quality. Fig. 12 is a diagram showing an example of an area requiring additional restoration predicted by the confidence map.
[0079] Next, the learning method of the device of the present invention is described. For example, YouTube-VOS can be used for learning the device of the present invention. This dataset consists of 2,932 training videos and 420 test videos. For a given image, two masks are generated, as shown in Fig. 13. The first is a mask that segments an object located in the image (Fig. 13 (a)). This mask is used to train the model to generate an object by learning the object's optical flow. The second is an outpainting mask (Fig. 13 (b)), which is used to learn the function of filling in the outer region. The ratio of inpainting masks to outpainting masks is 8:2.
[0080] < Loss function >
[0081] Four loss functions were applied to the present invention.
[0082] (1) loss function
[0083] First, the restored image with the outer region created Wow original video Compute the pixel-level difference between Use a loss function. The loss function is defined as follows:
[0084]
[0085] (2) loss function
[0086] Secondly, Confidence loss function to learn the extent of loss . The confidence loss function is For areas with a high degree of loss, it is used as an indicator for performing the creation task again.
[0087]
[0088] (3) loss function
[0089] Third, we use an adversarial loss, which has been proven effective in generating high-quality images. In this invention, we use a discriminator based on T-PATCHGAN, which is effective in reconstructing features of temporally overlapping regions. (See Ya-Liang Chang, Zhe Yu Liu, Kuan-Ying Lee, Winston Hsu, Free-form Video Inpainting with 3D Gated Convolution and Temporal PatchGAN, InarXiv:1904.10247.) The discriminator's loss function is defined as follows.
[0090]
[0091] The adversarial loss function of the generator is defined as follows:
[0092]
[0093] (4) loss function
[0094] Finally, loss of optical flow consistency . This loss function is used to generate flow in the masked area. Flow calculated from GT It serves to reduce the difference between the two and is defined as follows.
[0095]
[0096] The final loss function from the above four loss functions is defined as follows:
[0097]
[0098] In one embodiment, the Adam Optimizer was used for optimization, , was set to . For the following experiment, a total of 500K iterations were trained, and the initial learning rate was set to 0.0001.
[0099] < Experimental Results >
[0100] The quantitative and qualitative evaluation results of the method of the present invention are as follows.
[0101] (1) Quantitative evaluation
[0102] For quantitative evaluation, two dramas were selected, eight video clips were extracted from each video, and the average PSNR score was calculated. Figure 14 shows the results of the quantitative evaluation experiment of the present invention for the first drama, and Figure 15 shows the results of the quantitative evaluation experiment of the present invention for the second drama.
[0103] Baseline is the result of applying E2-FGVI, R is the result of applying an iterative mask, and CR is the result of applying an additional confidence mask. Both the iterative mask and the confidence mask contribute to improved performance.
[0104] (2) Qualitative evaluation
[0105] 1) Outer surface creation result of the present invention
[0106] Figure 16 shows a captured screen of an original video with a 4:3 aspect ratio and a video expanded to a 16:9 aspect ratio using the method of the present invention. As can be seen in the captured screen, the background is naturally generated even when played back as a video.
[0107] 2) Ability to maintain peripheral objects
[0108] Figure 17 is a diagram showing the results of applying the method of the present invention to an image expanded to a 16:9 ratio when an object exists on the periphery. As can be seen in Figure 17, it can be confirmed that even if a person is located on the periphery in an image expanded to 16:9, the image is restored.
[0109] 3) Performance comparison with conventional methods
[0110] The following videos were generated using existing SOTA models and played back using the method of the present invention. Figure 18 shows the results when applying the conventional method and the method of the present invention. The comparison with the conventional method confirmed that the method of the present invention outperforms the existing methods.
[0111] 4) Results applied to animation
[0112] Figure 19 shows the resulting screen when the method of the present invention is applied to animation. Figure 19 shows a screen (bottom) converted to a 16:9 ratio by applying the method of the present invention to a 4:3 ratio animation screen (top). As can be seen in Figure 19, the method of the present invention works well even for animation.
[0113] The above description is merely an example of the technical idea of the present embodiment, and those skilled in the art to which the present embodiment pertains may make various modifications and variations without departing from the essential characteristics of the present embodiment. Therefore, the present embodiments are not intended to limit the technical idea of the present embodiment, but to explain it, and the scope of the technical idea of the present embodiment is not limited by these embodiments. The protection scope of the present embodiment should be interpreted by the claims below, and all technical ideas within a scope equivalent thereto should be interpreted as being included in the scope of the rights of the present embodiment.
[0114] (Explanation of symbols)
[0115] 110 Repetitive Outer Extension Wrapper,
[0116] 120 mask filling deep network,
[0117] 130 Confidence Predictor.
[0118] CROSS-REFERENCE TO RELATED APPLICATION
[0119] This patent application claims priority to Korean Patent Application No. 10-2023-0146436, filed on October 30, 2023, the entire contents of which are incorporated herein by reference.
Claims
1. A mask filling deep network that receives an image and mask as input and generates a restored image. A confidence predictor that predicts the confidence of a restored image, and An iterative outer expansion wrapper that controls filling in the outer region of a video by repeating N operations in the mask filling deep network and the confidence predictor until an image of a second aspect ratio is generated, given an input video of a first aspect ratio. A device for generating the outer part of a video for changing the aspect ratio of a video having a .
2. In paragraph 1, The above iterative outer expansion wrapper inputs image[i] and mask[i] to the mask filling deep network and runs the deep network, The above mask filling deep network outputs a restored image [i] with the outer edges filled using the input image [i] and mask [i], The above confidence predictor outputs a confidence map[i] indicating the confidence of the filled area, Mask[i] is generated by applying thresholding and taking the pixel-wise sum of maski and the confidence map[i-1] generated in the previous operation. For image [i], the restored image [i-1] generated from the previous operation is used. A device that generates the outer part of a video to change the aspect ratio of the video.
3. In the second paragraph, the repetitive outer extension wrapper, Controlling the size of mask i to be increased gradually between the first screen ratio and the second screen ratio while performing N iterative operations. A device that generates the outer part of a video to change the aspect ratio of the video.
4. In paragraph 3, The output of the confidence predictor is a confidence map of the same size as the image size of the second aspect ratio, and the value of each pixel of the confidence map has a value of 0 or 1 based on whether the confidence value is greater than or equal to a predetermined value. A device that generates the outer part of a video to change the aspect ratio of the video.
5. In the second paragraph, the repetitive outer extension wrapper, The restored image is output by performing N iterative operations and is output as the final output image. A device that generates the outer part of a video to change the aspect ratio of the video.
6. In any one of paragraphs 1 to 5, the mask filling deep network, A flow completion unit that extracts and completes optical flow between adjacent frames from input images. CNN encoder that encodes the features of the input image, A feature propagation unit that calculates what to fill in the outer part of the image using the optical flow features generated in the flow completion unit and the context features encoded in the CNN encoder. A temporal transformer that performs the task of generating an outer region by combining information from nearby and non-near neighbors, and A CNN decoder that outputs a reconstructed image with the edges filled in from the information synthesized from the temporal transformer. Equipped with, The output of the above temporal transformer is input to the confidence predictor, A device that generates the outer part of a video to change the aspect ratio of the video.
7. In paragraph 6, The outer region generation device of the above video is trained using a first mask used to learn the function of generating an object by learning the optical flow of the object, and a second mask used to learn the function of filling the outer region. The first mask is a segmentation mask for objects located in the image, and the second mask is an outpainting mask. A device that generates the outer part of a video to change the aspect ratio of the video.
8. The first step is to input the input video of the first aspect ratio and the initial mask, Mask1, into the mask filling deep network using the iterative outer expansion wrapper, A second step in which a mask filling deep network outputs a restored image with the outer edges filled using the input image and mask, and a confidence predictor outputs a confidence map indicating the confidence of the filled area; The third step is to generate mask[i] by performing a pixel-wise sum of maski and the confidence map[i-1] generated in the previous operation and applying thresholding, and input mask[i] and the restored image[i-1] generated in the previous operation into the mask filling deep network. A fourth step of controlling the above-mentioned iterative outer expansion wrapper to repeat the second and third steps a predetermined number of times to output a restored image of the second aspect ratio size output from the mask filling deep network as the final output image. A method for generating an outer portion of a video for changing the aspect ratio of the video.
9. In paragraph 8, The above iterative outer expansion wrapper controls the N iterative operations while gradually increasing the size of mask i between the first aspect ratio and the second aspect ratio. How to create an outline of a video to change the aspect ratio of the video.
10. In paragraph 9, The output of the above confidence predictor is a confidence map of the same size as the image size of the second aspect ratio, and the value of each pixel of the confidence map has a value of 0 or 1 based on whether the confidence value is greater than or equal to a predetermined value. How to create an outline of a video to change the aspect ratio of the video.
11. In any one of paragraphs 8 to 10, the mask filling deep network, A flow completion unit that extracts and completes optical flow between adjacent frames from input images. CNN encoder that encodes the features of the input image, A feature propagation unit that calculates what to fill in the outer part of the image using the optical flow features generated in the flow completion unit and the context features encoded in the CNN encoder. A temporal transformer that performs the task of generating an outer region by combining information from nearby and non-near neighbors, and A CNN decoder that outputs a reconstructed image with the edges filled in from the information synthesized from the temporal transformer. Equipped with, The output of the above temporal transformer is input to the confidence predictor, How to create an outline of a video to change the aspect ratio of the video.
12. In the 11th paragraph, the method for generating the outer part of the video, It further comprises a step of learning using a first mask used to learn the function of generating an object by learning the optical flow of the object, and a second mask used to learn the function of filling in the outer region. The first mask is a segmentation mask for objects located in the image, and the second mask is an outpainting mask. How to create an outline of a video to change the aspect ratio of the video.
Citation Information
Patent Citations
Time-Synchronous Off-Chain Combined Payment System and Payment Method Based on Unique Contract Signature of Unique Identification Information
KR1020240050742A
Fender panel mounting structure of vehicle body
KR1020250005604A
Image out-painting appratus and method on deep-learning
KR102298175B1