Image Resynthesis Using Forward Warping, Gap Discriminator, and Coordinate-Based Inpainting

By using forward distortion and gap filling modules in image resynthesis, the problem of insufficient accuracy of image reconstitution caused by back distortion in the prior art is solved, and a higher quality new view generation is achieved.

CN112823375BActive Publication Date: 2025-07-22SAMSUNG ELECTRONICS CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN201980066712.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-04-29
Filing Date
2019-11-07
Publication Date
2025-07-22
Estimated Expiration
2039-11-07

AI Technical Summary

Technical Problem

The prior art uses backward distortion in image resynthesis, resulting in insufficient accuracy of image resynthesis, especially when scene modeling is difficult, it is difficult to generate high-quality new views.

Method used

The forward twist module is used to predict the position of the source image pixel and the target image, and the gap generated by the twist is filled with the gap filling module, and parallel training is carried out in combination with the gap discriminator network to improve the accuracy of image resynthesis.

Benefits of technology

Through the combination of forward distortion and gap filling modules, the accuracy of image resynthesis is significantly improved, resulting in a more realistic new view.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112823375B_ABST
    Figure CN112823375B_ABST
Patent Text Reader

Abstract

The present invention relates to image processing, and more particularly, to image resynthesis for synthesizing a new view of a person or an object based on an input image to solve tasks such as predicting views of a person or an object from a new viewpoint and a new pose. The technical result is to improve the accuracy of image resynthesis based on at least one input image. An image resynthesis system, a system for training an inpainting module to be used in the image resynthesis system, an image resynthesis method, a computer program product, and a computer-readable medium are provided. The image resynthesis system includes a source image input module, a forward warping module, and an inpainting module. The forward warping module is configured to predict a corresponding position in a target image for each source image pixel, and the forward warping module is configured to predict a forward warping field aligned with the source image. The inpainting module is configured to fill gaps generated from the application of the forward warping module. The image resynthesis method includes the steps of: inputting a source image, predicting a corresponding position in a target image for each source image pixel, wherein a forward warping field aligned with the source image is predicted; predicting a binary mask of gaps generated from the forward warping, generating a texture image by predicting a pair of coordinates in the source image for each pixel in the texture image, filling the gaps based on the binary mask of the gaps, and mapping the complete texture back to the new pose using backward warping.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to image processing, and more particularly, to image resynthesis for synthesizing new views of a person or object based on an input image using machine learning techniques. Background Art

[0002] Recently, there has been a growing interest in learning-based image resynthesis. In this context, the task of machine learning is to learn to synthesize new views of, for example, a particular type of person or object based on one or more input images of the person or object. In the extreme case, only one input view is available. In this sense, the new view corresponds to a new camera position and / or a new body pose of the person. In image resynthesis, the quality of the target view is measured, and the quality of the intermediate representation that is usually implicitly or explicitly associated with a model of the scene (e.g., 3D reconstruction) is not concerned. Directly optimizing the target view quality usually means a higher quality of the target view, especially when scene modeling is difficult.

[0003] Several trends have been found. First, dealing with the hard prediction problem accompanying image resynthesis requires a deep convolutional network (ConvNet) (see

[15] ). Second, many prior art solutions avoid directly predicting pixel values from high-dimensional non-convolutional representations. Instead, most architectures resort to some kind of warping within the ConvNet (see, for example, [5, 30, 20, 3, 23]). It is well known that in many cases, the prior art uses backward warping

[13] , in which, for each pixel in the target image, the position where the pixel in the source image will be copied is predicted. After the warping process, post-processing such as brightness correction (see [5]) or a post-processing network is usually performed.

[0004] Now several approaches to the problems related to the objective technical problems to be solved by the present invention will be discussed.

[0005] Warping-based resynthesis. There is a strong interest in using deep convolutional networks to generate realistic images (see, for example, [6]). When generating new images by changing the geometry and appearance of the input image, it has been shown that using a warping module greatly enhances the quality of the resynthesized image (see, for example, [5, 30]). In this case, the warping module is based on the differentiable (backward) grid sampler layer that was first introduced as part of a spatial transformer network (STN) (see, for example,

[13] ).

[0006] Adversarial image inpainting. There are also prior art solutions for image inpainting based on deep convolutional networks. Special variants of convolutional architectures adapted to gaps present in the input data include the Shepard convolutional neural network (see, e.g.,

[21] ), the sparse invariant convolutional network (see, e.g.,

[25] ), the network with partial convolution (see, e.g.,

[17] ), the network with gated convolution (see, e.g.,

[28] ). The latter variant is also used in the method presented in this disclosure.

[0007] Since the inpainting task requires conditional synthesis of image content, prior art inpainting methods heavily rely on variants of generative adversarial learning (see, e.g., [6]). Specifically, prior art suggests using discriminator pairs that focus on distinguishing between real and fake examples at two different scales (see, e.g., [16, 11, 28, 4]), where one of the scales can correspond to individual patches (similar to the patch GAN idea from

[12] ). Here, a novel discriminator is introduced, where the novel discriminator has a similar architecture to some local discriminators and patch GANs, yet discriminates between two different classes of pixels (known pixels in the fake image vs. unknown pixels also in the fake image).

[0008] Face frontalization. Prior art solutions focused on image resynthesis (such as generating new views and / or changing the pose of a 3D object based on a single input photographic image or multiple input photographic images) use images of faces as the main domain. The frontalized face view can be used as a normalized representation to simplify face recognition and enhance its quality. Several prior art solutions use backward samplers for this task. For example, the HF-PIM system, which can be considered the most typical example of such a method, predicts the cylindrical texture map and the backward warping field required to transform the texture map into a frontalized face view. The warped result is then corrected by another network. Many other methods currently considered efficient (such as CAPG-GAN (see, e.g., [9]), LB-GAN (see, e.g., [2]), CPF (see, e.g.,

[26] ), FF-GAN (see, e.g.,

[27] )) are based on encoder-decoder networks that directly perform the desired transformation by representing the image in a low-dimensional latent space. Additionally, the resynthesis network is typically trained in a GAN setting to make the output face look realistic and prevent various artifacts. Many of these methods employ additional information, such as landmarks (see, e.g., [9, 29]), local patches (see, e.g.,

[10] ), 3D deformation model (3DMM, see, e.g., [1]) estimation (see, e.g.,

[27] ). Then, such additional information can be used to regulate the resynthesis process or formulate additional losses by measuring how well the synthesized image conforms to the available additional information.

[0009] Full body recomposition based on warping. In the prior art, warping is utilized to synthesize a new view of a person in the case of a single input view (see, for example, [24, 23, 19]). This method also utilizes dense-pose parameterization within the network (see, for example, [8]) to present the target pose of the person on the recomposed image.

[0010] It should be noted that all the above types of prior art methods for image recomposition have certain drawbacks, and the present invention aims to eliminate or at least mitigate at least some of the drawbacks of the prior art. Specifically, the drawbacks of the available prior art solutions relate to the use of backward warping in image recomposition, in which, for each pixel in the target image, the position where the pixel in the source image will be copied is predicted. Summary of the Invention

[0011] Technical Problem

[0012] An object of the present invention is to provide a new method for image recomposition that eliminates or at least mitigates all or at least some of the above drawbacks of the existing prior art solutions.

[0013] The technical result achieved by the present invention is to improve the accuracy of image recomposition for synthesizing a new view of a person or an object based on at least one input image.

[0014] Technical Solution

[0015] In one aspect, this object is achieved by an image recomposition system, wherein the image recomposition system includes: a source image input module; a forward warping module configured to predict the corresponding position in the target image for each source image pixel, wherein the forward warping module is configured to predict a forward warping field aligned with the source image; and a gap filling module configured to fill the gaps generated from the application of the forward warping module.

[0016] In an embodiment, the gap filling module may further include: a warping error correction module configured to correct the forward warping error in the target image.

[0017] Advantageous Effects

[0018] The image recomposition system may further include a texture transfer architecture, wherein the texture transfer architecture is configured to: predict the warping fields for the source image and the target image; map the source image to the texture space via forward warping, restore the texture space to a complete texture; and map the complete texture back to the new pose using backward warping.

[0019] The image recomposition system may further include a texture extraction module, wherein the texture extraction module is configured to extract textures from the source image. At least the forward warping module and the gap filling module may be implemented as deep convolutional neural networks.

[0020] In an embodiment, the gap filling module may include a gap fixer, wherein the gap fixer includes: a coordinate assignment module configured to assign a pair of texture coordinates (u, v) to each pixel p = (x, y) of the input image according to a fixed predefined texture mapping, so as to provide a two-channel mapping of the x value and the y value in the texture coordinate system; a texture map completion module configured to provide a complete texture map, wherein for each texture pixel (u, v), the corresponding image pixel (x[u, v], y[u, v]) is known; a final texture generation module configured to generate a final texture by mapping the image value from the position (x[u, v], y[u, v]) to the texture at the position (u, v), so as to provide a complete color final texture; and a final texture remapping module configured to remap the final texture to a new view by providing a different mapping from the image pixel coordinates to the texture coordinates.

[0021] At least one of the deep convolutional neural networks in the deep convolutional neural network may be trained using a true / false discriminator configured to distinguish between a ground truth image and a repaired image. The image recomposition system may further include an image correction module configured to correct defects in the output image.

[0022] In another aspect, there is provided a system for training a gap filling module, wherein the gap filling module is configured to fill gaps as part of image recomposition, the system is configured to perform parallel joint training of the gap filling module with a gap discriminator network, and the gap discriminator network is trained to predict a binary mask of the gap, and the gap filling module is trained to minimize the accuracy of the gap discriminator network.

[0023] In yet another aspect, the present invention relates to an image recomposition method, including the following steps: inputting a source image; predicting a corresponding position in a target image for each source image pixel, wherein a forward warping field aligned with the source image is predicted; predicting a binary mask of a gap generated from the forward warping, generating a texture image by predicting a pair of coordinates in the source image for each pixel in the texture image, filling the gap based on the binary mask of the gap; and mapping the complete texture back to a new pose using backward warping.

[0024] In an embodiment, the step of filling the gap may include the following steps: Assigning a pair of texture coordinates (u, v) to each pixel p = (x, y) of the input image according to a fixed predefined texture mapping to provide a two-channel mapping of the x value and the y value in the texture coordinate system; Providing a complete texture map, wherein for each texture pixel (u, v), the corresponding image pixel (x[u, v], y[u, v]) is known; Generating a final texture by mapping the image value from the position (x[u, v], y[u, v]) to the texture at the position (u, v) to provide a complete color final texture; Remapping the final texture to a new view by providing a different mapping from the image pixel coordinates to the texture coordinates.

[0025] In yet another aspect, the present invention provides a method for training a gap filling module, wherein the gap filling module is configured to fill a gap as part of image resynthesis, the method comprising: Performing parallel joint training of the gap filling module with a gap discriminator network, while the gap discriminator network is trained to predict a binary mask of the gap, and the gap filling module is trained to minimize the accuracy of the gap discriminator network.

[0026] In yet another aspect, there is provided a computer program product comprising computer program code, wherein the computer program code, when executed by one or more processors, causes the one or more processors to implement the method of the second foregoing aspect.

[0027] In yet another aspect, there is provided a non-transitory computer-readable medium storing a computer program product according to the foregoing aspect.

[0028] By reading and understanding the specification provided below, those skilled in the art will understand that the claimed invention may also take other forms. Various method steps and system components can be implemented by means of hardware, software, firmware, or any suitable combination thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] After the above-provided summary of the invention, a detailed description of the inventive concept is provided below by way of example and with reference to the accompanying drawings, wherein the drawings are provided only as illustrations and are not intended to limit the scope of the claimed invention or to determine its essential features. In the drawings:

[0030] Figure 1 Shows the difference between forward warping and backward warping explained in terms of the face straightening task;

[0031] Figure 2 Shows a machine learning process for repair using a gap discriminator according to an embodiment of the present invention;

[0032] Figure 3Shows the process of frontal face normalization via forward warping according to an embodiment of the present invention;

[0033] Figure 4 Shows the texture transfer architecture according to an embodiment of the present invention;

[0034] Figure 5 Shows the process of texture completion for novel pose resynthesis according to an embodiment of the present invention;

[0035] Figure 6 Shows the process of full - body resynthesis using coordinate - based texture inpainting according to an embodiment of the present invention.

[0036] Figure 7 Shows a flowchart of an image resynthesis method according to an embodiment of the present invention. Detailed Description

[0037] This detailed description is provided to facilitate understanding of the essence of the present invention. It should be noted that the description relates to exemplary embodiments of the present invention, and those skilled in the art can envision other modifications, variations, and equivalent substitutions in the described subject matter by carefully reading this description with reference to the accompanying drawings. All such obvious modifications, variations, and equivalents are considered to be covered by the scope of the claimed invention. The reference numerals or symbols provided in this detailed description and in the claims are not intended to limit or define the scope of the claimed invention in any way.

[0038] The present invention proposes a new method for image resynthesis based on at least one input image. The systems and methods of the present invention are based on various neural networks that can be trained on various datasets, such as deep convolutional neural networks. It will be apparent to those skilled in the art that the implementation of the present invention is not limited to the neural networks specifically described herein, but other types of networks suitable for a given task can be used within the context of the present invention. The neural networks suitable for implementing the present invention can be realized by materials and technical means known to those skilled in the art (such as, but not limited to, one or more processors, general - purpose or special - purpose computers, graphics processing units (GPUs), etc., controlled by one or more computer programs, computer program elements, program codes, etc.) in order to implement the inventive methods described below.

[0039] First, the claimed invention will now be described in terms of one or more machine - learning models based on deep convolutional neural networks that are pre - trained or trained to perform a specified process, where the specified process results in image resynthesis for synthesizing a new view of a person or object based on at least one input image.

[0040] In summary, the proposed method is based on two contributions compared to the prior art. As a first contribution, a resynthesis architecture based on forward warping is proposed. Specifically, the warping process employed within the warping stage of the prior art method outlined above is redesigned by replacing the backward warping widely used in the prior art with a forward warping provided by a module that predicts the corresponding position in the target image for each source image pixel. The inventors have found that since the forward warp field is aligned with the source image, predicting the forward warp field from the source image is an easier task. This is in contrast to the backward warp field, which is spatially aligned with the target image and not spatially aligned with the source image. The presence of spatial alignment between the source image and the forward warp field makes the prediction mapping easier to learn for the convolutional architecture.

[0041] However, the result of the forward warping contains gaps that need to be filled. Most prior art solutions use adversarial architectures to solve the problem of gap inpainting. Therefore, the second contribution of the proposed invention is a novel gap discriminator dedicated to the inpainting task. The gap discriminator is trained only on "fake" (i.e., inpainted) images, and no "real" images are required. For each fake image, the gap discriminator is trained to predict a binary mask of the gaps provided to the inpainting network. Therefore, training on the gap discriminator causes the inpainting network to fill the gaps in a way that makes the gaps imperceptible.

[0042] The two proposed contributions are not independent but complement each other, forming a new resynthesis method that has been evaluated by the inventors for several tasks such as face straightening and full body resynthesis.

[0043] A new method for full body resynthesis is also proposed. In this method, the so-called DensePose method is used to estimate the body texture coordinates. A deep convolutional network is used to complete the texture. The deep convolutional network can be used to predict the color of unknown pixels even. Optionally, a deep network is used that predicts the coordinate pairs in the source image for each pixel in the texture image (coordinate-based inpainting). The latter scheme (coordinate-based inpainting) gives a clearer texture. The completed texture is used to generate a new view of the whole body, considering the body surface coordinates for each foreground pixel in the target image. Optionally, another deep network can be used to generate the final target image by taking as input the generated image with superimposed texture and some other images.

[0044] According to a first aspect, the present invention provides an image resynthesis system 100, comprising:

[0045] Source image input module 110;

[0046] forward warping module 120;

[0047] Gap filling module 130 .

[0048] The gap filling module 130 further includes a distortion error correction module 131 configured to correct the forward distortion error in the target image. The forward distortion module 120 is configured to predict the corresponding position in the target image for each source image pixel, and the forward distortion module is configured to predict the forward distortion field aligned with the source image. The gap filling module 130 is configured to fill the gaps generated from the application of the forward distortion module 120.

[0049] In an embodiment, the image recomposition system 100 further includes a texture transfer architecture 150, wherein the texture transfer architecture 150 is configured to perform the following operations: predicting the distortion fields for the source image and the target image; mapping the source image to the texture space via forward distortion, restoring the texture space to a complete texture; and mapping the complete texture back to the new pose using backward distortion.

[0050] In an exemplary embodiment, the image recomposition system 100 further includes a texture extraction module 160 configured to extract the texture from the source image. At least the forward distortion module 120 and the gap filling module 130 can be implemented as deep convolutional neural networks. At least one of these deep convolutional networks is trained using a true / false discriminator configured to distinguish between a ground truth image and an inpainted image.

[0051] In an embodiment, the gap filling module 130 includes a gap restorer 132, wherein the gap restorer 132 can at least consist of the following items in sequence:

[0052] A coordinate assignment module 1321 configured to assign a pair of texture coordinates (u, v) to each pixel p = (x, y) of the input image according to a fixed predefined texture mapping, so as to provide a two-channel mapping of the x value and the y value in the texture coordinate system;

[0053] A texture map completion module 1322 configured to provide a complete texture map, wherein for each texture pixel (u, v), the corresponding image pixel (x[u, v], y[u, v]) is known;

[0054] A final texture generation module 1323 configured to generate a final texture by mapping the image value from the position (x[u, v], y[u, v]) to the texture at the position (u, v), so as to provide a complete color final texture;

[0055] A final texture remapping module 1342 configured to remap the final texture to a new view by providing a different mapping from the image pixel coordinates to the texture coordinates.

[0056] In an embodiment, the image recomposition system 100 further includes an image correction module 170 configured to correct output image defects.

[0057] In another aspect of the present invention, a system 200 for training the gap filling module 130 is provided. The system 200 is configured to perform parallel joint training of the gap filling module with the gap discriminator network 210, and the gap discriminator network 210 is trained to predict a binary mask of the gap, and the gap filling module 130 is trained to minimize the accuracy of the gap discriminator network 210.

[0058] Referring Figure 7 , in yet another aspect, the present invention relates to an image recomposition method 300, wherein the method includes the following steps:

[0059] Input a source image (S310);

[0060] For each source image pixel, predict the corresponding position in the target image (S320), wherein a forward warping field aligned with the source image is predicted;

[0061] Predict a binary mask of the gap generated from the forward warping (S330),

[0062] Generate a texture image by predicting a pair of coordinates in the source image for each pixel in the texture image, and fill the gap based on the binary mask of the gap (S340); and

[0063] Use backward warping to map the complete texture back to the new pose (S350).

[0064] In an exemplary embodiment, the step of filling the gap (S340) includes the following steps:

[0065] Assign a pair of texture coordinates (u, v) to each pixel p = (x, y) of the input image according to a fixed predefined texture mapping (S341), so as to provide a two-channel mapping of the x value and the y value in the texture coordinate system;

[0066] Provide a complete texture map (S342), wherein for each texture pixel (u, v), the corresponding image pixel (x[u, v], y[u, v]) is known;

[0067] Generate a final texture by mapping the image value from the position (x[u, v], y[u, v]) to the texture at the position (u, v) (S343), so as to provide a complete color final texture;

[0068] Remap the final texture to the new view by providing a different mapping from the image pixel coordinates to the texture coordinates (S344).

[0069] There is also provided a computer program product 400 including computer program code 410, wherein the computer program code 410, when executed by one or more processors, causes the one or more processors to implement the method according to the previous aspect. The computer program product 400 can be stored on a non-transitory computer-readable medium 500.

[0070] Now referring to Figure 1 , the difference between forward warping and backward warping as explained in the face alignment task is shown. In both scenarios, a warping field (bottom; hue = direction, saturation = magnitude) is predicted from the input image (top), and the warping is applied (right). In the case of forward warping, the input image and the predicted field are aligned (e.g., the movement of the nose tip is predicted at the position of the nose tip). Conversely, in the case of backward warping, the input image and the warping field are not aligned. The forward warping method in the context of the present invention will now be described in more detail.

[0071] It will be readily understood by those skilled in the art that the methods described below are suitable for being performed by a deep convolutional neural network, wherein the deep convolutional neural network can implement the elements of the image resynthesis system 100 and the steps of the image resynthesis method 300 as described above. The detailed description of the methods provided below with reference to the mathematical operations and the relationships between various data elements may depend on the corresponding functions rather than the specific elements of the system 100 or the method 300 outlined above, and in this case, those skilled in the art can readily deduce on the one hand the relationships between the system elements and / or method steps, and on the other hand the corresponding functions mentioned below, without strictly limiting the scope of the various ways of implementing the functions by the specific associations between each function and the corresponding system element and / or method step. In the context of implementing the method of the present invention for image resynthesis as described in detail below, the system elements and / or method steps implemented by the deep convolutional neural network are intended to be exemplary and non-limiting.

[0072] Resynthesis by Forward Warping

[0073] Let x be the source image and let y be the target image, and let x[p, q] denote the image entry (sample) at the integer position (p, q) (which can be, for example, an RGB value). Let w[p, q] = (u[p, q], v[p, q]) be the warping field. Generally, the warping field will be predicted from x by a convolutional network f θ , where θ is a vector of some learnable parameters trained on a specific dataset.

[0074] The standard method for resynthesis of a warped image uses the warping to warp the source image x to the target image y:

[0075] ybw [p, q] = x[p + u[p, q], q + v[p, q]] (1)

[0076] Among them, the sampling at the fractional positions is defined bilinearly. More formally, the result of the backward warping is defined as:

[0077]

[0078] Among them, the bilinear kernel K is defined as follows:

[0079] K(k, l, m, n) = max(1 - |m - k|, 0)max(1 - |n - l|, 0), (3)

[0080] So that for each (P, q), the summation in (2) is performed for i = {|p + u[p, q]|, |p + u[p, q]|} and j = {[q + v[p, q]], |q + v[p, q]|}.

[0081] The backward warping method was initially implemented for depth image recognition (see, for example,

[13] ), and later was widely used for depth image resynthesis (see, for example, [5, 30, 20, 3, 23]), becoming a standard layer within the deep learning package. It has been found that for resynthesis tasks with significant geometric transformations, compared with architectures that resynthesize using only convolutional layers, using the backward warping layer provides significant improvements in quality and generalization ability (see, for example, [3]).

[0082] However, the backward warping is limited by the lack of alignment between the source image and the warping field. In fact, as can be seen from the expression (1) provided above, the vector (u[p, q], v[p, q]) predicted by the network f θ for the pixel (p, q) defines the motion of the object part that was initially projected onto the pixel (p + u[p, q], q + v[p, q]). For example, let's consider the face frontalization task, where, in the case of an input image containing a non-frontal face, it is expected that the network predicts the frontalization warping field. Assume that the position (p, q) in the initial image corresponds to the tip of the nose, while for the frontalized face, the same position corresponds to the center of the right cheek. When the backward warping is used for resynthesis, the network f θ the prediction for the position (p, q) must include the frontalization motion of the center of the right cheek. At the same time, the receptive field of the output network unit at (p, q) in the input image corresponds to the tip of the nose. Therefore, the network must predict the motion of the cheek while observing the appearance of the block centered on the nose (see Figure 1 ). When the frontalization motion is small, such misalignment can be handled by a deep enough convolutional architecture with a receptive field large enough. However, as the motion becomes larger, for convolutional architectures, such a mapping becomes increasingly difficult to learn.

[0083] Therefore, forward warping performed by a forward warping module is used instead of backward warping in the resynthesis architecture according to the present invention. The forward warping operation is defined such that the following equation approximately holds for the output image y fw approximately holds:

[0084] y fw [p + u[p, q], q + v[p, q]] ≈ x[p, q]. (4)

[0085] Therefore, in the case of forward warping, the warping vector at pixel [p, q] defines the motion of that pixel. To implement forward warping, a bilinear kernel is used to rasterize source pixels onto the target image in the following manner. First, all contributions from all pixels are aggregated into an aggregator map a using a convolutional kernel:

[0086]

[0087] At the same time, the total weight of all contributions for each pixel is accumulated in a separate aggregator w:

[0088]

[0089] Finally, the value at the pixel is defined by normalization:

[0090] y fw [i, j] = a[i, j] / (w[i, j] + ∈), (7)

[0091] where the small constant ∈ = 10 -10 prevents numerical instability. Formally, for each target position (i, j), the summations in (5) and (6) are performed over all source pixels (p, q). However, since for each source pixel (p, q), the bilinear kernel K(·, ·, p + u[p, q], q + v[p, q]) takes non - zero values only at four positions in the target image, the above summations can be efficiently computed using a single pass over the pixels of the source image. Note that a similar technique is used in partial convolution (see, e.g.,

[17] ). Since the operations (5) - (7) are piece - wise differentiable with respect to both the input image x and the warping field (u, v), gradients can be backpropagated through the forward warping operation during convolutional network training.

[0092] The main advantage of forward warping over backward warping is that, in the case of forward warping, the input image and the predicted warping field are aligned because the prediction of the network at pixel (p, q) now corresponds to the 2D motion of the object part projected onto (p, q) in the input image. In the example of uprighting shown above, the convolutional network has to predict the uprighting motion of the nose tip based on the receptive field centered at the nose tip. For the convolutional network, this mapping is easier to learn than in the case of backward warping, and this effect has been proven by experiments.

[0093] However, disadvantageously, in most cases, the output y of the forward warping operation fw contains multiple empty pixels to which no source pixels are mapped. The binary mask of non-empty pixels is denoted as m, i.e., m[i, j] = [w[i, j] > 0]. Then, the following repair phase is required to fill such gaps.

[0094] Figure 2 Illustrated is the process of training a neural network to perform gap repair using a gap discriminator. The present inventors trained a repair network to fill gaps in an input image (where the known pixels are specified by a mask) while minimizing the reconstruction loss with respect to the "ground truth". In parallel, a segmentation network (also referred to here as a gap discriminator) was trained to predict the mask from the result of the filling operation while minimizing the mask prediction loss. The repair network was adversarially trained against the gap discriminator network by maximizing the mask prediction loss, which results in the filled part in the reconstructed image being indistinguishable from the original part.

[0095] Now, the process of "repairing" the gaps generated from the previous stage of forward warping will be described in more detail by way of illustration rather than limitation.

[0096] Repair using a gap discriminator

[0097] An image completion function g with learnable parameters φ φ maps the image y fw and the mask m to the complete (repaired) image y inp :

[0098] y inp = g φ (y fw , m). (8)

[0099] It has been experimentally proven that the operation of using a deep network with gated convolution to handle the repair task effectively provides a good architecture for repairing gaps generated from warping in the process of image resynthesis. Regardless of g φRegardless of the architecture, the choice of the loss function used for its learning plays a crucial role. Most commonly, training φ is done in a supervised setting, which involves providing a dataset of full images, devising a random process for occluding parts of those images, and training a network to invert that random process. Then, minimization of the subsequent loss is performed at training time:

[0100]

[0101] where i iterates over the training examples, and denotes the full image. The norm in expression (9) can be chosen as the L1 norm (i.e., the sum of the absolute differences in each coordinate) or as a more sophisticated perceptual loss that is not based on the differences between pixels but on the differences between high-level image feature representations extracted from a pre-trained convolutional neural network (see, for example,

[14] ).

[0102] When the empty pixels form large continuous gaps, due to the inherent multimodality of the task, pixelwise learning results or perceptual losses are usually suboptimal and lack reasonable large-scale structure. The use of adversarial learning (see, for example, [6]) gives a significant boost in this case. Adversarial learning performs parallel training of a separate classification network d ψ with the network g φ For the training of d ψ the training objective is to distinguish between the inpainted image and the original (undamaged) image:

[0103]

[0104] Then, separate terms are used to enhance the training objective for g φ where the separate terms measure the probability that the discriminator classifies the inpainted image as a real image:

[0105]

[0106] Existing state-of-the-art methods for adversarial inpainting propose using two discriminators, both based on the same principle but focusing on different parts of the image. One of these discriminators (called the global discriminator) focuses on the whole image, while the other of these discriminators (the local discriminator) focuses on the most important parts (such as the immediate vicinity of the gap or the central part of the face) (see, for example, [4]).

[0107] The present invention proposes to use different kinds of discriminators (here called gap discriminators) for gap repair tasks. The inventors have found that humans tend to judge the success of a repair operation by their ability (or lack thereof) to identify gap regions in the repaired image. Interestingly, for such a judgment, humans do not need to know any kind of "ground truth". To mimic this idea, the gap discriminator h ξ is trained to predict a mask m from the repaired image by minimizing a weighted cross-entropy loss for binary segmentation:

[0108]

[0109] where ⊙ denotes the element-wise product (summed over all pixels), and |m| denotes the number of non-zero pixels in the mask m. As the gap discriminator is being trained, the repair network is trained to confound the gap discriminator by maximizing the same cross-entropy loss (12) (thus playing a zero-sum game). The new loss can be used together with the "traditional" adversarial loss (11) and any other losses. The proposed new loss is applicable to any repair / completion problem without having to incorporate forward warping.

[0110] Learning with incomplete ground truth. In some cases (such as texture repair tasks), a complete ground truth image is not available. Instead, each ground truth image has a binary mask of known pixels which must be different from the input mask m i (otherwise, the training process may converge to a trivial identity solution for the repair network). In this case, the losses characterized in the above expressions (9)-(11) are adapted such that y i and are accordingly replaced by and Interestingly, the new adversarial loss does not consider the complete ground truth image. Thus, even when a complete ground truth image is not available, the loss characterized in the above expression (12) can still be applied without modification (to both gap discriminator training and as the loss for repair network training).

[0111] Figure 3 Shows an example of face frontalization via forward warping according to at least one embodiment of the present invention. In this example, a visual evaluation was performed on an algorithm trained on 80% of the random samples from the Multi-PIE dataset, based on two randomly selected subjects from the validation section. Each input photo ( Figure 3 the first row in) independently generates a warped image via a warp field regressor ( Figure 3in the second row), and then a repaired image filled with gaps and corrected for warping errors is generated by a fixer (third row).

[0112] Figure 4 Shows an example of a texture transfer architecture according to at least one embodiment of the present invention. The texture transfer architecture predicts the warping fields for both the source image and the target image. Then, by mapping the source image into the texture space via forward warping, the source image is restored to a complete texture, and then it is mapped back to the new pose using backward warping, and then the result is corrected. In Figure 4 where and are forward warping and backward warping respectively, while and WF are the dense pose warping fields of the predicted value and the ground truth.

[0113] Figure 5 Shows an example of texture completion for the task of new pose resynthesis according to at least one embodiment of the present invention. The texture is extracted from the input image of a person ( Figure 5 the first column in). Then the texture is repaired using a deep convolutional network. Figure 5 The third column of Figure 5 shows the result of repair using a network trained without a gap discriminator. Adding a gap discriminator ( Figure 5 the fourth column in) produces a sharper and more reasonable result in the repaired area. Then the resulting texture is overlaid on the image of the person in the new pose (the fifth and sixth columns in Figure 5 respectively).

[0114] Figure 6 Shows an example of full-body resynthesis using coordinate-based texture repair. The input image A is used to generate the texture image B. Coordinate-based repair is applied to predict the coordinates of the pixels in the source image for each texture image. The result is shown in image C, where the color of the texture pixels is sampled from the source image at the specified coordinates. The image of the person in the new pose (the target image) is synthesized by obtaining the specified texture coordinates of the pixels of the target image and transferring the color from the texture (image D). Finally, a separate correction depth network transforms this image into a new image (image E). Image F shows the ground truth image of the person in the new pose.

[0115] End-to-end training

[0116] Since both the forward warping and the inpainting network are end-to-end differentiable (i.e., the partial derivatives of any loss function with respect to the parameters of all layers, including the layers before the forward warping module and the layers of the inpainting network, can be computed using the backpropagation process), the (forward warping and inpainting) joint system can be trained in an end-to-end manner while applying a gap discriminator that attempts to predict the gap positions resulting from the forward warping process to the combined network.

[0117] Coordinate-based inpainting

[0118] The goal of coordinate-based inpainting is to complete the texture of, for example, a person depicted in an image based on a portion of the texture extracted from a source image. More specifically, starting from the source image, the following steps are performed:

[0119] 1. A pre-trained deep neural network is run to assign a pair of texture coordinates (u, v) to each pixel p = (x, y) of the input image according to a fixed predefined texture mapping. As a result, a subset of texture pixels is assigned pixel coordinates, thereby producing a two-channel mapping with x-values and y-values in a texture coordinate system with a large number of texture pixels that do not know the mapping.

[0120] 2. As a next step, a second deep convolutional neural network h with learnable parameters μ is run such that the x and y mappings are completed, thereby producing a complete texture map, where for each texture pixel (u, v), the corresponding image pixel (x[u, v], y[u, v]) is known.

[0121] 3. The final texture is obtained by taking the image values (e.g., in the red, green, and blue channels) at the positions (x[u, v], y[u, v]) and placing them on the texture at the positions (u, v), thereby producing a complete color texture.

[0122] 4. Once the complete texture is obtained, it is used to texture new views of a person in different poses, where a different mapping from image pixel coordinates to texture coordinates is provided for the new views.

[0123] When creating the texture using a first image of the pair of texture coordinates (u, v), the parameters μ can be further optimized such that the order of steps 1 - 4 outlined above closely matches a second image in the pair of texture coordinates (u, v). Any standard loss (e.g., pixel-wise, perceptual) can be used to measure the closeness. A gap discriminator can be added to the training of the texture completion network.

[0124] Finally, a separate refinement network can be used to transform the re-textured image obtained in the order of steps 1 - 4 in order to improve the visual quality. The refinement network can be trained either separately or jointly with the texture completion network.

[0125] Next, a specific example of a practical implementation of the inventive method is provided by way of illustration and not limitation.The inventive method based on forward warping followed by an inpainting network trained with a gap discriminator is applied to various tasks with different levels of complexity.

[0126] Facial correction

[0127] As a first task, we consider a face rotation method that aims to warp a non-frontally oriented face into a normalized face while preserving identity, facial expression, and illumination. We train and evaluate this method on the Multi-PIE dataset, e.g. as described in [7], which is a dataset of over 750,000 upper body images of 337 people imaged at four conferences with varying (and known) views, illumination conditions, and facial expressions. We use a U-Net-shaped (see e.g.

[22] ) architecture (N convolutional layers, N for the forward warp, and an hourglass-shaped architecture for the inpainting network).

[0128] Face and upper body rotation

[0129] For face rotation, the proposed method is trained and evaluated on the Multi-PIE dataset (see, e.g., [7]). For each subject, 15 views are captured simultaneously by multiple cameras, of which 13 cameras are placed around the subject in the same horizontal plane at regular intervals of 15°, ranging from -90° to 90°, and 2 cameras are at an elevated level. Each multi-view set is captured under 19 different illumination conditions, up to 4 sessions and 4 facial expressions. In the experiments, only the 13 cameras placed around the subject in the same horizontal plane are used. For the upper body experiments, the original images are used, while for the face experiments, the MTCNN face detector is used to find the face bounding box and crop it with a gap of 10 pixels. 128×128 is the standard resolution for the experiments, and all images are finally resized to this resolution before passing through the learning algorithm. This rotation is considered to be the most important special case of the rotation task of the experiments.

[0130] The proposed image resynthesis pipeline consists of two parts: a warp field regressor implemented as a forward warping module and a repairer implemented as a gap filling module. The warp field regressor is a convolutional network (ConvNet) with trainable parameters ωf ω , where the convolutional network follows the U-Net architecture (see, e.g.,

[22] ). Given an input image (and two additional grid arrays encoding pixel rows and pixel columns), the ConvNet produces an offset field w encoded by two 2D arrays δ [p, q] = (u δ[p, q], v δ [p, q]). This field is later transformed into a forward warping field by a simple addition w[p, q] = (u[p, q], v[p, q]) = (p + u δ [p, q], q + v δ [p, q]) and passed through a forward grid sampler. In the described case, w δ [p, q] encodes the motion of the pixel (p, q) on the input image. However, note that if a backward sampler is further applied, the same structure can potentially be used to regress the backward warping field.

[0131] In the second part, the inpainter is a network g with learnable parameters φ φ , where the network g φ is also based on a U-Net architecture in which all convolutions are replaced by gated convolutions (although there are no skip connections). These are attention layers first proposed in

[28] to effectively handle difficult inpainting tasks. The gated convolution used is as defined in

[28] :

[0132] gated = conu(I, W g ),

[0133] feature = conu(I, W f ),

[0134] output = ELU(feature) · σ(gated), (13)

[0135] where is the input image, is the weight tensor, and σ and ELU are the sigmoid and exponential linear unit activation functions respectively. The inpainter receives a warped image with gaps, a gap mask, and a grid tensor encoding the positions of the pixels, and predicts the inpainted image.

[0136] The model is trained in a generative adversarial network (GAN) framework, and two discriminator networks are added. The first discriminator (real / fake discriminator) aims to distinguish the ground truth output image from the inpainted image produced by the generative inpainting network. The real / fake discriminator d ψ can be organized as a stack of plain convolutions and strided convolutions, mainly following the architecture of the VGG-16 feature extractor part, and then followed by average pooling and sigmoid. The resulting number indicates the predicted probability that the image is a "real" image. The second discriminator is a gap discriminator h that aims to recover the gap mask from the inpainted image by solving a segmentation problem ξ . Conversely, the GAN generator tries to "fool" the gap discriminator by producing an image with inpainted regions indistinguishable from the non-inpainted regions.

[0137] As described above, end-to-end learning of the pipeline is a difficult task that requires careful balancing among various loss components. The loss value Lgenerator for the generated ConvNet is optimized as follows, where the generated ConvNet includes a warping field regressor followed by a restorer:

[0138] L gencrator (ω, φ) = L warping (ω) + L inpaintcr (ω, φ) +

[0139] α adv L adv (ω, φ) + α gap L gap (ω, φ), (14)

[0140] where L warping (ω) penalizes the warped image and the warping field, and L inpainter (ω, φ) only penalizes the restored image, L adv (ω, φ) and L gap (ω, φ) are generator penalties corresponding to adversarial learning using the first discriminator (real / fake discriminator) and the second discriminator (gap discriminator), respectively. Thus, these components are decomposed into the following basic loss functions:

[0141]

[0142]

[0143]

[0144]

[0145] where is the warped image and the non-gap mask obtained by the forward sampler, is the forward warping field, and x i is the i-th input sample. For clarity, the grid that is the input to the warping field regressor and the restorer is omitted here and below.

[0146]

[0147]

[0148]

[0149] where v is the identity feature extractor. Light-CNN-29 pre-trained on the MS-Celeb-1M dataset is used as the source of identity-invariant embeddings. During training, the weights of v are fixed.

[0150] L adv (ω, φ) follows expression (11), and L gap (ω, φ) is similarly defined as:

[0151]

[0152] Together with the generator, the two discriminators are updated by the above losses (10) and (12).

[0153] Return Figure 3 , showing the performance of the algorithm trained on a subset of the Multi-PIE with 80% of the random samples and evaluated on the remaining 20% of the data. Figure 3 The results shown in

[0154] Body texture estimation and pose transfer

[0155] For the texture transfer task, forward warping and gap discriminator techniques are adopted. The DeepFashion dataset (see, for example,

[18] ) is used to demonstrate the performance of the constructed model, where the model restores the complete texture of the human body, and the complete texture of the human body can be projected onto a body of any shape and pose to generate a target image.

[0156] The model performs texture transfer in four steps:

[0157] 1. Map the initial image to the texture space and detect its missing parts

[0158] 2. Repair the missing parts to restore the complete texture

[0159] 3. Project the restored texture onto the new body pose

[0160] 4. Correct the resulting image to eliminate the defects that appear after texture reprojection.

[0161] Although separate loss functions are applied to the outputs of each module, all the modules that perform these steps can be trained simultaneously in an end-to-end manner. The scheme of the model can be viewed in Figure 5 .

[0162] Twenty-four textures of different body parts located on a single RGB image are used to predict the texture image coordinates of each pixel on the source image, in order to map the input image into the texture space by generating their warping fields. The warping fields for both the source image and the target image are generated by the same network with a similar UNet architecture. Since the warping field establishes the correspondence between the initial image and the texture space coordinates, the texture image can be generated from the photo using forward warping, and the human image can be reconstructed from the texture using backward warping with the same warping field. To train the warping field generator, a ground truth uv renderer generated by a dense pose model (see, for example, [8]) is used, and the L1 (sum of absolute differences) loss is used to penalize the resulting warping field:

[0163] L warp (x source , WF, ω) = ||w ω (x source ) - WF||1, (18)

[0164] where x source is the source image, WF is the ground truth warping field, and w ω is the warping field generator.

[0165] After mapping the source image into the texture space, due to self-occlusion on the source image, the resulting texture image has many missing parts (gaps). Then, these missing parts are repaired by a gap filling module, where the gap filling module is the second part of the trained model.

[0166] The trained model adopts a gated an hourglass architecture (see, for example,

[28] ) and a gap discriminator to generate reasonable textures as well as the L1 loss. The loss function of the repairer is as follows:[[]]

[0167] L inpainter (t source , t target , φ, ξ) =,[[]]

[0168] ||g φ (t source ) - t source ||1 +,[[]]

[0169] ||g φ (t source ) - t target ||1 +,[[]]

[0170] L gap (φ, ξ), (19)

[0171] Here, t source and t targetare the source texture and the target texture, respectively, g φ is the restorer and calculates L as in expression (12) gap , and for the loss function 10, update the weights of the gap discriminator.

[0172] Once the restored texture is generated, the restored texture can be reprojected back onto any body encoded by its warping field. By generating the warping field for the target image by the warping field predictor, an image of the source person in the target pose can be generated by backward warping the texture using the target warping field.

[0173] Although reasonable texture reconstruction can be obtained in this way, the resulting warped images may have many defects caused by differences observed when different texture parts are joined and some smaller body regions missing in the texture space. It is difficult to solve these problems in the texture space. However, they can be easily handled in the original image space. For this purpose, an image correction module implemented as a correction network can be used to eliminate these defects. The output of the correction network is the final result of the model. The VGG loss between the final result and the true target image and the true / false discriminator that attempts to distinguish between images from the dataset and the generated images are calculated:

[0174] H(t source , x target , φ, ω) = g φ (t source ),

[0175]

[0176] L refine (t source , x target , φ, ω, ψ) =,[[]]

[0177] VGG(H(t source , x target , φ, ω), x target ) -,[[]]

[0178] log(1 - s ψ (H(t source , x target , φ, ω))), (20)

[0180] where, is backward warping, VGG is the VGG loss, i.e., the 12 distance between the features extracted by the VGG-16 network, and s ψ is the true / false discriminator, and its loss function is expressed as follows:

[0181] L real / fake=-log(s ψ (H(t source , x target , φ, ω))) (21).

[0182] One or more machine learning models based on a deep convolutional neural network pre-trained or being trained to perform an image resynthesis task have now been described. The image resynthesis system 100 of the present invention implementing the method according to the present invention can be specifically characterized as including:

[0183] A source image input module 110; a forward warping module 120 configured to predict a corresponding position in the target image for each source image pixel, wherein the forward warping module 120 is configured to predict a forward warping field aligned with the source image; a gap filling module 130 including a gap discriminator 210 and a gap restorer 132, wherein the gap discriminator 210 is configured to predict a binary mask of the gap generated from the forward warping, and the gap restorer 132 is configured to generate a texture image by predicting a pair of coordinates in the source image for each pixel in the texture image, and fill the gap based on the binary mask of the gap; and a target image output module 180.

[0184] The forward warping module 120 may further include a warping field regressor 121 for generating a forward warped image. The gap filling module 130 as described above may further include a warping error correction module 131 configured to correct the forward warping error in the target image.

[0185] According to at least one embodiment, the system of the present invention may further include a texture transfer architecture 150 configured to perform the following operations: predict warping fields for the source image and the target image; map the source image into the texture space via forward warping, restore the texture space to a complete texture; and map the complete texture back to the new pose using backward warping.

[0186] The system may further include a texture extraction module 160 configured to extract texture from a source image. According to the present invention, at least the forward warping module 120 and the inpainting module 130 may be implemented as deep convolutional neural networks. At least one of the deep convolutional networks in the deep convolutional network may be trained using a true / false discriminator configured to distinguish between a ground truth image and an inpainted image. The gap discriminator 210 of the system of the present invention may be trained in the form of a separate classification network to distinguish between an inpainted image and an original image by minimizing a weighted cross-entropy loss for binary segmentation to predict a mask m from the inpainted image. The gap inpainter 132 of the system of the present invention may further include: a coordinate assignment module 1321 configured to assign a pair of texture coordinates (u, v) to each pixel p = (x, y) of an input image according to a fixed predefined texture mapping to provide a two-channel mapping of x values and y values in a texture coordinate system; a texture map completion module 1322 configured to provide a complete texture map, wherein for each texture pixel (u, v), the corresponding image pixel (x[u, v], y[u, v]) is known; a final texture generation module 1323 configured to generate a final texture by mapping image values from the position (x[u, v], y[u, v]) to the texture at the position (u, v) to provide a complete color final texture; and a final texture remapping module 1324 configured to remap the final texture to a new view by providing a different mapping from image pixel coordinates to texture coordinates.

[0187] In at least one embodiment, the image resynthesis system may include an image correction module 170, wherein the image correction module 170 is configured to correct output image defects caused by differences observed when different texture portions are joined.

[0188] It will be apparent to those skilled in the art that the above modules of the system of the present invention may be implemented by various means well known in the art, such as various software, hardware, firmware, etc. For example, various combinations of hardware and software may be contemplated for performing the above functions and / or processes, and these combinations will be apparent to those skilled in the art upon careful study of the description provided above. The claimed invention is not limited to any particular form of implementation or combination as described above, but may be implemented in various forms according to the particular image resynthesis task to be solved.

[0189] The following is a detailed description of specific exemplary embodiments of the present invention, where the specific exemplary embodiments are intended to illustrate rather than limit the materials and technical means for implementing the components of the corresponding image processing system and the steps of the image processing method, their functional properties, the relationships between them, and the operating modes of the image processing system and method of the present invention. After carefully reading the above-provided description with reference to the accompanying drawings, other embodiments falling within the scope of the present invention may become apparent to those skilled in the art, and all such apparent modifications, variations, and / or equivalent replacements are considered to be covered by the scope of the present invention. The order of the steps of the method of the present invention recited in the claims need not define the actual order in which the method steps are intended to be performed, and the specific method steps may be performed substantially simultaneously, one after another, or in any suitable order, unless the context of the present disclosure specifically defines and / or stipulates otherwise. Although not specified in the claims or elsewhere in the application materials, the specific method steps may also be performed once or any other suitable number of times.

[0190] It should also be noted that the present invention may take other forms compared to the forms described above, and in applicable cases, specific components, modules, elements, functions may be implemented as software, hardware, firmware, integrated circuits, FPGAs, etc. The claimed invention or at least its specific components, modules, elements, or steps may be implemented by a computer program stored on a computer-readable medium, where the program, when running on a general-purpose computer, GPU, multifunctional device, or any suitable image processing device, causes the device to perform some or all of the steps of the claimed method and / or control at least some components of the claimed image resynthesis system such that they operate in the above-described manner. Examples of computer-readable media suitable for storing the computer program or its code, instructions, or computer program elements or modules may include any kind of non-transitory computer-readable medium that is obvious to those skilled in the art.

[0191] All non-patent prior art references [1]-

[30] cited and discussed in this document and listed below are hereby incorporated by reference into the present disclosure where applicable.

[0192] By reading the above description with reference to the accompanying drawings, further aspects of the present invention may occur to those skilled in the art. Those skilled in the art will recognize that other embodiments of the present invention are feasible, and the details of the present invention may be modified in many aspects without departing from the inventive concept. Therefore, the drawings and the description are considered to be illustrative rather than restrictive in nature. The scope of the claimed invention is determined only by the claims.

[0193] Citation List

[0194] [1] V. Blanz and T. Vetter. Face recognition based on fitting a 3d morphable model. T-PAMI, 25(9):1063 - 1074, 2003.2

[0195] [2] J. Cao, Y. Hu, B. Yu, R. He, and Z. Sun. Load balanced gans for multi-view face image synthesis. arXiv preprint arXiv:1802.07447, 2018.2

[0196] [3] J. Cao, Y. Hu, H. Zhang, R. He, and Z. Sun. Learning a high fidelity pose - invariant model for high - resolution face frontalization. arXiv preprint arXiv:1806.08472, 2018.1,3

[0197] [4] J. Deng, S. Cheng, N. Xue, Y. Zhou, and S. Zafeiriou. Uv - gan: adversarial facial uv map completion for pose - invariant face recognition. In Proc. CVPR, pages 7093 - 7102, 2018.1,2,4

[0198] [5] Y. Ganin, D. Kononenko, D. Sungatullina, and V. Lempitsky. Deepwarp: Photorealistic image resynthesis for gaze manipulation. In European Conference on Computer Vision, pages 311 - 326. Springer, 2016.1,2,3

[0199] [6]I. Goodfellow, J. Pouget - Abadie, M. Mirza, B. Xu, D. Warde - Farley, S. Ozair, A. Courville, and Y. Bengio. Generative adversarial nets. In Advances in neural information processing systems, pages 2672 - 2680, 2014. 2, 4

[0200] [7]R. Gross, I. Matthews, J. Cohn, T. Kanade, and S. Baker. Multi - pie. Image and Vision Computing, 28(5):807 - 813, 2010. 5

[0201] [8]R. A. Guler, N. Neverova, and I. Kokkinos. DensePose: Dense human pose estimation in the wild. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2018. 2, 8

[0202] [9]Y. Hu, X. Wu, B. Yu, R. He, and Z. Sun. Pose - guided photo - realistic face rotation. In Proc. CVPR, 2018. 2

[0203]

[10] R. Huang, S. Zhang, T. Li, R. He, et al. Beyond face rotation: Global and local perception gan for photorealistic and identity preserving frontal view synthesis. arXiv preprint arXiv:1704.04086, 2017. 2

[0204]

[11] S. Iizuka, E. Simo-Serra, and H. Ishikawa. Globally and locally consistent image completion. ACM Transactions on Graphics (TOG), 36(4):107, 2017. 1, 2

[0205]

[12] P. Isola, J. Zhu, T. Zhou, and A. A. Efros. Image-to-image translation with conditional adversarial networks. In Proc. CVPR, pages 5967 - 5976, 2017. 2

[0206]

[13] M. Jaderberg, K. Simonyan, A. Zisserman, and K. Kavukcuoglu. Spatial transformer networks. In Proc. NIPS, pages 2017 - 2025, 2015. 1, 2, 3

[0207]

[14] J. Johnson, A. Alahi, and L. Fei-Fei. Perceptual losses for real-time style transfer and super-resolution. In Proc. ECCV, pages 694 - 711, 2016. 4

[0208]

[15] Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel. Backpropagation applied to handwritten zip code recognition. Neural computation, 1(4):541 - 551, 1989. 1

[0209]

[16] Y. Li, S. Liu, J. Yang, and M.-H. Yang. Generative face completion. In Proc. CVPR, volume 1, page 3, 2017. 1, 2

[0210]

[17] G. Liu, F. A. Reda, K. J. Shih, T.-C. Wang, A. Tao, and B. Catanzaro. Image inpainting for irregular holes using partial convolutions. arXiv preprint arXiv:1804.07723, 2018. 1, 2, 3

[0211]

[18] Z. Liu, P. Luo, S. Qiu, X. Wang, and X. Tang. Deep fashion: Powering robust clothes recognition and retrieval with rich annotations. In Proc. CVPR, pages 1096 - 1104, 2016. 6

[0212]

[19] N. Neverova, R. A. Guler, and I. Kokkinos. Dense pose transfer. In The European Conference on Computer Vision (ECCV), September 2018. 2

[0213]

[20] E. Park, J. Yang, E. Yumer, D. Ceylan, and A. C. Berg. Transformation - grounded image generation network for novel 3d view synthesis. In Proc. CVPR, pages 702 - 711. IEEE, 2017. 1, 3

[0214]

[21] J. S. Ren, L. Xu, Q. Yan, and W. Sun. Shepard convolutional neural networks. In Proc. NIPS, pages 901 - 909, 2015. 1, 2

[0215]

[22] O. Ronneberger, P. Fischer, and T. Brox. U - net: Convolutional networks for biomedical image segmentation. In Proc. MICCAI, pages 234 - 241. Springer, 2015. 5

[0216]

[23] A. Siarohin, E. Sangineto, S. Lathuilire, and N. Sebe. Generative adversarial networks for pose-based human image generation. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2018. 1, 2, 3

[0217]

[24] S. Tulyakov, M.-Y. Liu, X. Yang, and J. Kautz. Moco-gan: Decomposing motion and content for video generation. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2018. 2

[0218]

[25] J. Uhrig, N. Schneider, L. Schneider, U. Franke, T. Brox, and A. Geiger. Sparsity invariant cnns. In International Conference on 3D Vision (3DV), pages 11 - 20. IEEE, 2017. 1, 2

[0219]

[26] J. Yim, H. Jung, B. Yoo, C. Choi, D. Park, and J. Kim. Rotating your face using multi-task deep neural network. In Proc. CVPR, pages 676 - 684, 2015. 2

[0220]

[27] X. Yin, X. Yu, K. Sohn, X. Liu, and M. Chandraker. Towards large-pose face frontalization in the wild. In Proc. ICCV, pages 1 - 10, 2017. 2

[0221]

[28] J.Yu, Z.Lin, J.Yang, X.Shen, X.Lu, and T.S.Huang. Free-form image inpainting with gated convolution. arXiv preprint arXiv:1806.03589, 2018. 1, 2, 5, 6, 8

[0222]

[29] J.Zhao, L.Xiong, P.K.Jayashree, J.Li, F.Zhao, Z.Wang, P.S.Pranata, P.S.Shen, S.Yan, and J.Feng. Dual-agent gans for photorealistic and identity preserving profile face synthesis. In Proc.NIPS, pages 66 - 76, 2017. 2

[0223]

[30] T.Zhou, S.Tulsiani, W.Sun, J.Malik, and A.A.Efros. View synthesis by appearance flow. In Proc.ECCV, pages 2016. 1, 2, 3.

Claims

1. An image recomposition system, comprising: A source image input module; A forward warping module configured to predict a corresponding position in a target image for each source image pixel, wherein the forward warping module is configured to predict a forward warping field aligned with the source image; And A gap filling module configured to predict a binary mask of a gap resulting from the application of the forward warping module and fill the gap based on the binary mask by predicting a pair of coordinates in the source image for each pixel in a texture image to generate a texture image.

2. The image resynthesis system according to claim 1, wherein, The gap filling module further includes a warping error correction module, wherein the warping error correction module is configured to correct a forward warping error in the target image.

3. The image resynthesis system according to claim 1 further includes: A texture transfer architecture configured to: Predict warping fields for the source image and the target image; Map the source image to a texture space via forward warping, Restore the texture space to a complete texture; and Map the complete texture back to a new pose using backward warping.

4. The image re-synthesis system according to claim 1 further comprises: A texture extraction module configured to extract texture from the source image.

5. The image resynthesis system according to claim 1, wherein, At least the forward warping module and the gap filling module are implemented as deep convolutional neural networks.

6. The image re-synthesis system according to claim 1, wherein, The gap filling module includes a gap repairer, wherein the gap repairer includes: A coordinate assignment module configured to assign a pair of texture coordinates (u, v) to each pixel p = (x, y) of an input image according to a fixed predefined texture mapping to provide a two-channel mapping of x values and y values in a texture coordinate system; A texture map completion module configured to provide a complete texture map, wherein for each texture pixel (u, v), the corresponding image pixel (x[u, v], y[u, v]) is known; A final texture generation module configured to generate a final texture by mapping an image value from a position (x[u, v], y[u, v]) to a texture at a position (u, v) to provide a complete color final texture; A final texture remapping module configured to remap the final texture to a new view by providing a different mapping from image pixel coordinates to texture coordinates.

7. The image re-synthesis system according to claim 5, wherein, At least one of the deep convolutional neural networks in the deep convolutional neural network is trained using a true / false discriminator configured to distinguish between a ground truth image and a repaired image.

8. The image resynthesis system according to claim 4 further includes: An image correction module configured to correct output image defects.

9. A system for training a gap filling module, wherein, The gap filling module is configured to fill gaps as part of image recomposition, and the system is configured to perform parallel joint training of the gap filling module with a gap discriminator network, and the gap discriminator network is trained to predict a binary mask of a gap, and the gap filling module is trained to minimize the accuracy of the gap discriminator network.

10. An image recomposition method, comprising the following steps: Input a source image; For each source image pixel, predict a corresponding position in a target image, wherein a forward warping field aligned with the source image is predicted; Predict a binary mask of a gap resulting from forward warping, A texture image is generated by predicting a pair of coordinates in the source image for each pixel in the texture image, and the gap is filled based on the binary mask of the gap; and The complete texture is mapped back to the new pose using backward warping.

11. The image resynthesis method according to claim 10,[[]] Among them, The step of filling the gap includes the following steps: Assign a pair of texture coordinates (u, v) to each pixel p = (x, y) of the input image according to a fixed predefined texture mapping, so as to provide a two-channel mapping of the x value and the y value in the texture coordinate system; Provide a complete texture map, wherein for each texture pixel (u, v), the corresponding image pixel (x[u, v], y[u, v]) is known; Generate the final texture by mapping the image value from the position (x[u, v], y[u, v]) to the texture at the position (u, v), so as to provide a complete color final texture; Remap the final texture to a new view by providing a different mapping from image pixel coordinates to texture coordinates.

12. A method for training a gap filling module, wherein, The gap filling module is configured to fill the gap as part of image resynthesis, and the method includes: performing parallel joint training of the gap filling module with the gap discriminator network, and the gap discriminator network is trained to predict the binary mask of the gap, and the gap filling module is trained to minimize the accuracy of the gap discriminator network.

13. A computer program product comprising computer program code, wherein, The computer program code, when executed by one or more processors, causes the one or more processors to implement the method according to any one of claims 10 or 11.

14. A non-transitory computer-readable medium storing a computer program product according to claim 13.

Citation Information

Patent Citations

  • Locating features in warped images

    US20180060688A1