System and method for operating two-dimensional (2D) image of three-dimensional (3D) object
Patent Information
- Application Number
- JP2022141485
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-10-13
- Filing Date
- 2022-09-06
- Publication Date
- 2025-07-10
- Estimated Expiration
- 2042-09-06
AI Technical Summary
Existing image manipulation technologies struggle to create photorealistic 2D images of 3D objects, as current methods either lack the ability to handle large variations in pose changes or require high-quality 3D scans that are time-consuming and resource-intensive.
A system and method that manipulates 2D images of 3D objects by isolating and modifying physical properties such as shape, albedo, pose, and lighting using a neural network with sub-networks and a differentiable renderer, combined with a style-based GAN architecture for photorealistic rendering.
Enables efficient generation of photorealistic 2D images with variations in pose, expression, and lighting, reducing the need for high-quality 3D data and improving storage efficiency.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to image processing, and more specifically, to systems and methods for manipulating two-dimensional (2D) images of three-dimensional (3D) objects.
Background Art
[0002] Today, image manipulation of two-dimensional (2D) images of the human face has gained popularity in a wide range of applications. These include changing illumination conditions to create attractive portrait images, changing a person's identity to anonymize the image, and changing hairstyles in virtual try-on settings. In face image manipulation, carefully handling the details of the face image is extremely important for achieving photo-realism (realistic depiction like a photograph) in the synthesized face image. However, the human visual system can also be sensitive to very small artifacts in the synthesized face image. Typically, for a 2D image of an object to be photo-realistic, it should be equivalent to the 2D image of a real 3D object (e.g., a 3D face whose shape and reflection characteristics are within the range of typical human experience). Therefore, for a modified 2D face image to remain photo-realistic, it should correspond to real variations in the corresponding 3D object or its environment, such as variations in the apparent shape, expression, and appearance attributes of the corresponding 3D face. Existing prior art methods for photo-realistic automatic image manipulation of 2D face images are essentially 2D, and they perform the task of modifying one 2D face image to form another 2D face image without explicitly defining the corresponding 3D face for either image. This limits the ability of such 2D methods to synthesize large variations such as changes in the pose of a 2D image (e.g., from a frontal view to a side view).
[0003] On the other hand, existing methods for image manipulation, which are inherently 3D and trained from 3D scans of human faces, can perform 3D manipulations accurately, but they are limited in their ability to produce photorealistic 2D images of the resulting 3D-composite faces. For example, a 3D face model can be used to reconstruct a 3D face from a 2D image of a human face. The 3D model can be obtained from a training 3D dataset of human faces. The training 3D dataset may have limited variation in the attributes of the 2D images, which can result in limitations on facial expression and pose control. Furthermore, the 3D training dataset may not contain enough detail or sufficiently accurate information about the face to render a photorealistic 2D image. Such limitations can lead to a lack of realism in the reconstructed 2D face image. In some cases, high-quality 3D scans of human faces can be performed to increase the 3D data in the training 3D dataset. However, high-quality 3D scans may consume an excessive amount of time and storage resources.
[0004] Therefore, a technical solution is needed to reconstruct and manipulate 2D images that possess photorealism. More specifically, it is necessary to manipulate 2D images of 3D objects to achieve photorealism in an efficient and feasible manner. [Overview of the project]
[0005] Therefore, the object of this disclosure is to provide a system for manipulating 2D images of 3D objects. An image of a 3D object is a 2D representation of a 3D object that is affected by one or more physical properties of the 3D object. One or more physical properties of a 3D object include the 3D shape of the 3D object, the albedo of the 3D shape, the pose of the 3D object, and the lighting illuminating the 3D object. When there is a change in the value of each of these physical properties, the 2D image representation of the object is affected. For that purpose, the 2D image representation of the object can be manipulated based on the manipulation of each of the physical properties. In addition to, or instead of, prior to the manipulation, the 2D image can be masked by a mask image. The mask image can be automatically generated for the 2D image using a segmentation method such as a semantic segmentation network, landmark localization, or face analysis. The mask image segments regions. For example, it segments the foreground region to the background region in a 2D image corresponding to a 3D object by masking the foreground region to the background region. For example, a human face in a 2D image may be segmented by masking the background region in the 2D image using a corresponding face mask image of the human face, and further segmented into hair regions using a corresponding hair mask image of the hair. Such mask images can improve further processing of the 2D image representation of a 3D object.
[0006] Some embodiments are based on the understanding that one or more physical properties can be extracted from a 2D image representation of a 3D object. For example, some embodiments are based on the understanding that a neural network can be trained in one or more supervised and / or unsupervised modes to extract one or more physical properties of a 3D object from a training 2D image of the 3D object and to reconstruct the original training image from the extracted one or more physical properties of the object. One or more physical properties may include the 3D shape of the 3D input object, the albedo of the 3D shape of the 3D input object, the pose of the 3D input object, and the lighting illuminating the 3D input object. In some exemplary embodiments, the neural network is pre-trained in one or more supervised and / or unsupervised modes based on a dataset containing synthetically generated human face images. Such a pre-training process may enable the neural network to capture one or more prominent facial features from the synthetically generated human face images.
[0007] Some embodiments are based on the understanding that manipulating these extracted physical properties of a 3D object prior to image reconstruction can be used for realistic image manipulation by manipulating the properties of a 3D object and then rendering the manipulated 3D object as a 2D image, in contrast to direct image manipulation in the 2D domain of the image. Some embodiments are based on another understanding that a neural network may include multiple and / or different subnetworks. In some exemplary embodiments, multiple subnetworks extract different physical properties of a 3D object from a 2D image. The extraction of different physical properties by subnetworks allows for the separation of different physical properties in independent or distinct manners. Each independent separation of physical properties allows for the representation of different variations of one or more corresponding physical properties, which then contributes to achieving realism when reconstructing a 2D output image from a 2D input image of a 3D object. For that purpose, the separated physical properties can be rendered to become a 2D image.
[0008] Some embodiments are based on the recognition that the resulting rendered 2D image may not have the desired level of photorealism due to limitations in factors such as the rendering process, the representation of one or more physical properties, or the level and type of variation present in the training data. Some embodiments are based on yet another recognition that the photorealism of the rendered 2D image may be improved by 2D refinement of the rendered 2D image, which in some embodiments is done using a 2D refiner network that can be trained using an adversarial approach, such as the one used when training a generative adversarial network (GAN).
[0009] In some exemplary embodiments, each of one or more physical properties is isolated from the pixels of a 2D input image using multiple subnetworks by a trained neural network. Each of the multiple subnetworks isolates one or more of the one or more physical properties of a 3D object. The isolated one or more physical properties generated by the multiple subnetworks are combined using a differentiable renderer to form a 2D output image. For example, the differentiable renderer uses the isolated physical properties of lighting and pose of a 3D object to render a reconstructed 2D image, such as the 2D output image. Such isolation allows for the design of different specialized subnetwork architectures that can be adapted for different physical properties.
[0010] Some embodiments of using a differentiable renderer rather than a more traditional computer graphics renderer are based on the understanding that a combinatorial neural network, including one or more subnetworks and the differentiable renderer, can be thoroughly trained, and therefore errors in the rendered image can be propagated back through the combinatorial neural network to train one or more subnetworks.
[0011] In one embodiment, an albedo subnetwork may be trained to extract the albedo of a 3D object from pixels in a 2D input image. For this purpose, the albedo subnetwork includes a style-based generative adversarial network (GAN) architecture, and the network may be trained adversarially using a GAN framework. A style-based GAN may be trained using a set of photographic images of objects from a predetermined class. In some embodiments, the 3D object may be selected from a predetermined class of human faces.
[0012] Style-based GAN architectures may be advantageous for generating realistic 2D images.
[0013] Some embodiments are based on the understanding that style-based GANs are trained on a large dataset of 2D photographic images of objects from a predetermined class (e.g., a large database of face images) and then generate photorealistic 2D images of objects from a predetermined class (e.g., faces). Typically, style-based GANs are used to directly generate 2D images that are indistinguishable from photographic images from the training dataset, using adversarial training with a discriminator network. However, for a style-based GAN to generate a 3D object model using only 2D photorealistic image data, the output space of the style-based architecture may be part of the 3D model. The output space may not be the same space as the real-world dataset. Some embodiments are based on yet another understanding that style-based GANs can be used to generate the albedo of a 3D object, e.g., a 2D albedo map. Such a 2D albedo map may be applied to the 3D shape of the 3D object, which may further be input to a renderer to generate a composite 2D output image of the 3D object. However, generating synthetic 2D output images may require high-quality albedo maps to train style-based GANs by applying discriminators to that output space.
[0014] Some embodiments are based on the additional recognition that discriminator networks can be used for adversarial training of style-based GANs to generate photorealistic synthetic 2D output images without high-quality albedo maps. The discriminator network may be applied after using a differentiable renderer to render a 3D object model into a 2D synthetic image. The differentiable renderer that renders the 3D object model into a 2D synthetic image prevents the direct application of 2D albedo maps to the style-based GAN. Thus, the style-based GAN architecture may be used for the generated albedo maps for the 3D objects, while the discriminator network is used in the space of 2D images for the output at the end of the rendering process. The space of 2D images can be used to obtain large amounts of real-world training data (e.g., photographic images of faces).
[0015] In addition, or alternatively, some embodiments are based on the recognition that neural networks can be trained in an unsupervised manner. Unsupervised neural networks can extract different physical properties of 3D objects into latent spaces. Latent spaces can be used to reconstruct physical spaces for the corresponding physical properties. For example, albedo subnetworks and 3D shape subnetworks may have a GAN architecture that encodes pixels of a 2D image into their respective corresponding latent spaces for the object's albedo and 3D shape. These latent spaces can be used to generate physical spaces for the object's albedo and 3D shape. Latent and physical spaces enable dual representations of objects. Such dual representations allow for the manipulation of physical properties in either or both of the latent and physical spaces to achieve a photorealistic 2D image. In contrast, in some embodiments, illumination and pose subnetworks are encoded from pixels of a 2D image toward their corresponding physical spaces. Unsupervised training of neural networks for the extraction of different physical properties into latent spaces can eliminate the need for actual 3D data, which can improve the overall performance and storage efficiency of image processing systems.
[0016] In addition, or alternatively, the reconstructed 2D output image may include reconstructed hair on a human face in the 2D input image. In some embodiments, the reconstructed hair may be obtained from a hair model modeled by a hair manipulation algorithm. The hair manipulation algorithm manipulates the hair model separately from manipulating one or more physical properties of the human face. To that end, the hair manipulation algorithm modifies the representation of the hair in the 2D output image in response to detecting manipulation of the human face's pose from the 2D input image to the 2D output image. In some exemplary embodiments, the hair manipulation algorithm modifies the appearance of the hair to correspond to the modified pose of the human face in the 2D output image. Such a hair manipulation algorithm enables the output of a realistic composite image in which the hair corresponds to the hair in the 2D input image, modified to correspond to one or more manipulated physical properties, such as the modified pose.
[0017] In addition, or instead, an image processing system may be used to generate multiple realistic composite images from 2D images of a 3D object. These multiple realistic composite images may include variations in one or more physical properties of the 3D object, such as pose, facial expression, and lighting. These multiple realistic composite images may be used in different applications. These applications may include applications for generating cartoon or anime versions of a user, applications for experimenting with different hairstyles, camera-enabled robotic system applications, and face reconstruction applications. In some exemplary embodiments, multiple realistic composite images may be used as training data sample images for different systems (e.g., robotic systems, face reconstruction applications) to efficiently generate realistic training images while preventing the collection of sample images using cameras associated with the corresponding systems. For example, an existing dataset for training a different system may include facial images where variations in one or more physical properties, such as pose, facial expression, and lighting, are not the same as those when the different system is deployed. In such cases, multiple realistic composite images may be used to enhance the existing training dataset with images containing more examples for the different system, or to replace the existing training dataset with such images.
[0018] In addition, or instead, the image processing system may be used to generate an anonymized version of the input image. The anonymized version may be generated using an anonymization technique, such as a similarity-based metric. This may be used in applications where images, such as facial images, should be made public, but the system's users do not have permission to share photos or videos of one or more people in the input image.
[0019] Accordingly, one embodiment discloses an image processing system for manipulating two-dimensional (2D) images of a predetermined class of three-dimensional (3D) objects. The image processing system includes an input interface, memory, a processor, and an output interface. The input interface is configured to receive a 2D input image of a predetermined class of 3D input objects and one or more operation commands for manipulating one or more physical properties of the 3D input objects. The one or more physical properties include the 3D shape of the 3D input object, the albedo of the 3D shape of the 3D input object, the pose of the 3D input object, and the illumination illuminating the 3D input object. The memory is configured to store a neural network. The neural network is trained to reconstruct the 2D input image by separating one or more physical properties of the 3D input object from the pixels of the 2D input image using multiple subnetworks. The separated one or more physical properties generated by the multiple subnetworks are combined using a differentiable renderer to form a 2D output image. For each subset of one or more physical properties of a 3D input object, there is one subnetwork, and each subnetwork in a plurality of subnetworks is trained to extract the corresponding physical property from the pixels of a 2D input image into one or a combination of the latent space and / or physical space of the corresponding physical property. The albedo subnetwork, trained to extract the albedo of a 3D input object from the pixels of a 2D input image, has a style-based generative adversarial network (GAN) architecture trained using a set of photographic images of objects from a predetermined class. The processor is configured to present the 2D input image received from the input interface to the neural network and to manipulate the latent space, physical space, or both of the latent space and / or physical properties of one or more physical properties of the 3D input object according to one or more operation instructions.The operation is performed on one or more isolated physical properties generated by multiple subnetworks to modify the 3D input object before the differentiable renderer reconstructs the 2D output image. The output interface is configured to output a 2D output image.
[0020] Accordingly, another embodiment discloses a method for manipulating a two-dimensional (2D) image of a predetermined class of three-dimensional (3D) objects. The method includes receiving a 2D input image of a predetermined class of 3D input objects and one or more operation commands for manipulating one or more physical properties of the 3D input objects. The one or more physical properties include the 3D shape of the 3D input object, the albedo of the 3D input object, the pose of the 3D input object, and the illumination illuminating the 3D input object. The method includes presenting the 2D input image to a neural network and manipulating the latent space, physical space, or both of the one or more physical properties of the 3D input object according to one or more operation commands, the manipulation being performed on one or more isolated physical properties generated by a plurality of subnetworks to modify the 3D input object before a differentiable renderer reconstructs a 2D output image. The neural network is trained to reconstruct a 2D input image by using multiple subnetworks to separate the physical properties of a 3D input object from the pixels of a 2D input image, and then combining the separated physical properties generated by the subnetworks using a differentiable renderer to form a 2D output image. There is one subnetwork for each subset of one or more physical properties of the 3D input object. Each subnetwork in the multiple subnetworks is trained to extract the corresponding physical properties from the pixels of the 2D input image into one or a combination thereof of the latent space and / or physical space of the corresponding physical properties. The albedo subnetwork, trained to extract the albedo of the 3D input object from the pixels of the 2D input image, has a style-based generative adversarial network (GAN) architecture trained using a set of photographic images of objects from a predetermined class. The method further includes the step of outputting a 2D output image.
[0021] Further features and advantages will become more readily apparent from the following detailed description, which will be interpreted in conjunction with the attached drawings.
[0022] Brief Description of the Drawings This disclosure is further described in the following detailed description with reference to a plurality of drawings described as non-limiting examples of exemplary embodiments of the disclosure. Throughout several views of the drawings, the same reference numerals represent like parts. The drawings shown are not necessarily to scale; rather, emphasis is generally placed upon illustrating the principles of the embodiments disclosed herein.
Brief Description of the Drawings
[0023] [Figure 1] FIG. showing an exemplary representation depicting image manipulation of an image of a 3D object according to some embodiments of the present disclosure. [Figure 2] FIG. is a block diagram of a system for manipulating a two-dimensional (2D) image of a three-dimensional (3D) object of a predetermined class according to an exemplary embodiment of the present disclosure. [Figure 3] FIG. shows a framework of a neural network for manipulating a 2D image of a 3D object of a predetermined class according to an exemplary embodiment of the present disclosure. [Figure 4] FIG. is a flowchart showing a process for manipulating a 2D image of a 3D object of a predetermined class according to an exemplary embodiment of the present disclosure. [Figure 5] FIG. shows a framework corresponding to a hair manipulation algorithm according to an exemplary embodiment of the present disclosure. [Figure 6A] FIG. shows a method for manipulating a 2D image of a 3D object of a predetermined class according to an exemplary embodiment of the present disclosure. [Figure 6B] FIG. shows a method for manipulating a 2D image of a 3D object of a predetermined class according to an exemplary embodiment of the present disclosure. [Figure 7] FIG. is a block diagram of a system for manipulating a 2D image of a 3D object of a predetermined class according to an exemplary embodiment of the present disclosure. [Figure 8]A diagram showing a usage scenario for manipulating a 2D image of a 3D object of a predetermined class according to an exemplary embodiment of the present disclosure. [Figure 9] A diagram showing a usage scenario for manipulating a 2D image of a 3D object of a predetermined class according to another exemplary embodiment of the present disclosure. [Figure 10] A diagram showing a usage scenario for manipulating a 2D image of a 3D object of a predetermined class according to another exemplary embodiment of the present disclosure. [Figure 11] A diagram showing a usage scenario for manipulating a 2D image of a 3D object of a predetermined class according to yet another exemplary embodiment of the present disclosure.
[0024] The above drawings describe the embodiments disclosed herein. However, as described in the description, other embodiments are also conceivable. This disclosure is shown as an illustration and not a limitation of exemplary embodiments. Many other modifications and embodiments may be devised by those skilled in the art that fall within the scope and spirit of the principles of the embodiments disclosed herein.
Best Mode for Carrying Out the Invention
[0025] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure may be practiced without these specific details. In other instances, the apparatus and methods are shown only in block diagram form to avoid obscuring the present disclosure.
[0026] The terms “for example,” “such as,” and “etc.,” as used in this specification and claims, as well as the verbs “equip,” “have,” and “include,” and their other verbal forms, should each be interpreted as non-restrictive when used with a list of one or more components or other items; that is, the list should not be considered as excluding other additional components or items. The term “based on” means based at least partially. It should also be understood that the language and terminology used herein are for illustrative purposes only and should not be considered restrictive. Any headings used within this description are for convenience only and have no legal or restrictive effect.
[0027] Overview The proposed image processing system enables the reconstruction of 2D output images from 2D input images of a predetermined class of 3D objects, such as human faces. The objective of the image processing system is to manipulate the 2D input images for the reconstruction of the 2D output images. The 2D input images are manipulated by separating one or more physical properties of the 2D input image, such as the 3D shape of the 3D object, the albedo of the 3D object, the pose of the 3D object, and the lighting illuminating the 3D object. This is further illustrated in Figure 1.
[0028] Figure 1 shows an exemplary representation 100 of image manipulation of an image 102 of a 3D object according to some embodiments of the present disclosure. Image 102 is a 2D representation of a 3D object, such as a human face. Image 102 possesses one or more physical properties that can be extracted from the 2D representation of the human face. In some exemplary embodiments, one or more physical properties can be extracted from the pixel values of the 2D image 102. The extracted one or more physical properties can be explicitly represented as the estimated 3D shape 104A of the 3D input object, the estimated albedo 104B of the 3D input object, the estimated pose 104C of the 3D input object, and the estimated illumination 104D that illuminates the 3D input object. Each explicit representation of one or more physical properties, collectively referred to below as one or more physical properties 104A to 104D, enables independent manipulation of each of the one or more physical properties 104A to 104D. The operations here correspond to changes in the respective values of one or more corresponding physical properties 104A to 104D (e.g., pose angle or lighting parameter). When there is a change in the value of one or more physical properties, the 2D representation of image 102 is manipulated. In some exemplary embodiments, such independent manipulation of one or more physical properties 104A to 104D may be used to generate multiple realistic composite images from a 2D image of a 3D object. The multiple realistic composite images may correspond to variations in one or more physical properties 104A to 104D of the 3D object.
[0029] One or more physical properties 104A to 104D are combined to generate a 3D model of an object, which can be rendered to produce a 2D composite image 106A similar to the input image 102. One or more operations on the physical properties 104A to 104D correspond to changes in the 3D object model, which triggers corresponding operations in the 2D composite image 106A. Such operations, prior to rendering the 3D model as a reconfigured composite 2D image 106A, enable image manipulation of image 102 in a realistic manner. In some embodiments, the reconfigured composite 2D face model 106A is combined with a hair model 106B to produce a final output image, such as output image 108. Output image 108 is a photorealistic composite 2D image corresponding to the manipulated version of the input image 102.
[0030] Furthermore, each change in one or more physical properties 104A to 104D generates different operations on the image 102, such as lighting variations 110, facial expression variations 112, pose variations 114, shape variations 116, and texture variations 118. These different operations on the physical properties 104A to 104D can be used to synthesize photorealistic 2D images corresponding to variations in the input image 102 in terms of properties such as pose, facial expression, and lighting conditions.
[0031] Such image manipulation of the 2D representation of image 102, through the separation of one or more physical properties 104A to 104D, is performed by the system. This will be further explained with reference to Figure 2.
[0032] Figure 2 shows a block diagram of a system 200 for manipulating 2D images of a predetermined class of 3D objects, according to an exemplary embodiment of the present disclosure. The system 200 corresponds to an image processing system including an input interface 202, memory 204, a processor 206, and an output interface 208. The input interface 202 is configured to receive a 2D input image, such as image 102 of a 3D input object, i.e., a human face. The input interface 202 is also configured to receive one or more operation commands for manipulating one or more physical properties of the 3D input object, such as one or more physical properties 104A to 104D of the 3D input object, including the 3D shape of the 3D input object (e.g., shape 104A), the albedo of the 3D input object (e.g., albedo 104B), the pose of the 3D input object (e.g., pose 104C), and the illumination characteristics that illuminate the 3D input object (e.g., illumination 104D).
[0033] Memory 204 is configured to store the neural network 210. The neural network 210 is trained to reconstruct the image 102 by isolating one or more physical properties 104A to 104D. One or more physical properties 104A to 104D are isolated from pixels in the image 102. In some embodiments, multiple subnetworks are used to isolate one or more physical properties 104A to 104D from pixels. There is one subnetwork for each of one or more subsets of one or more physical properties 104A to 104D of the 3D object. Each subnetwork of one or more subnetworks is trained to extract the corresponding physical property from the pixels of the image 102 into one or a combination thereof of the latent space and / or physical space of the corresponding physical property. For example, the albedo subnetwork is trained to extract the albedo 104B of the 3D shape 104A of the 3D input object from the pixels of the image 102. In some embodiments, the albedo subnetwork may have a style-based generative adversarial network (GAN) architecture. In some exemplary embodiments, the style-based GAN may be trained using a set of photographic images of objects from a predetermined class.
[0034] Furthermore, each of the individual isolated physical properties is combined to form a 2D output image. In some embodiments, the individual isolated physical properties are combined using a differentiable renderer. The differentiable renderer uses the physical space of one or more physical properties 104A to 104D to render a 2D output image of image 102.
[0035] The processor 206 is configured to present the image 102 received from the input interface 202 to the neural network 210 while manipulating the latent space, physical space, or both of one or more physical properties of a 3D input object according to one or more operation instructions. The processor 206 is configured to manipulate physical properties separated by multiple subnetworks in order to modify the 3D input object so that the differentiable renderer reconstructs the 2D output image. The 2D output image is output via the output interface 208.
[0036] The separation of one or more physical properties by the corresponding subnetworks of the neural network 210 will be further explained with reference to Figure 3.
[0037] Figure 3 shows a framework 300 of a neural network 210 for manipulating 2D images of a predetermined class of 3D objects, according to an exemplary embodiment of the present disclosure. When image 102 is received, processor 206 generates a mask image 302 for image 102. In some exemplary embodiments, processor 206 may use a semantic segmentation network to estimate the mask image 302. The semantic segmentation network may generate a mask image 302 that includes some pixels representing the background (e.g., pixels with a value of zero) and the remaining pixels representing the foreground (e.g., pixels with a non-zero value, such as a value of 1.0). Image 102 is masked with mask image 302 such that the pixel values representing 3D objects in image 102 are preserved in the masked image 304, while the pixel values representing the background and (in the case of a face) hair in image 102 are masked by pixel values that are zero. Masking the pixel values of the segmented background (and hair) produces a masked image 304 containing pixel values that represent objects of a predetermined class (e.g., faces) in image 102. This masked image 304 is presented to the neural network 210.
[0038] The neural network 210 includes a plurality of subnetworks that separate one or more physical characteristics 104A to 104D of image 102. In some exemplary embodiments, the plurality of subnetworks correspond to an encoder architecture including a shape encoder 306A, an albedo encoder 306B, a pause encoder 306C, and an illumination encoder 306D (hereinafter referred to as encoders 306A to 306D), as shown in Figure 3. Each of the plurality of subnetworks, i.e., each of the encoders 306A to 306D, processes image 102 in such a way that each of the corresponding one or more physical characteristics 104A to 104D is extracted from the pixels of the masked image 304. For example, encoder 306A extracts shape 104A, encoder 306B extracts albedo 104B, encoder 306C extracts pause 104C, and encoder 306D extracts illumination 104D.
[0039] One or more physical properties 104A to 104D extracted from image 102 can be represented in one or a combination thereof of the latent space and / or physical space of the corresponding physical property. For example, shape encoder 306A extracts shape code 308A in the shape latent space, and albedo encoder 306B extracts albedo code 308B in the albedo latent space. In contrast, pause encoder 306C extracts pause representation 308C in the pause physical space, and encoder 306D extracts illumination representation 308D in the illumination physical space. Each of the shape code 308A and albedo code 308B represents a compressed state of the corresponding physical property, namely shape 104A and albedo 104B.
[0040] Furthermore, the shape code 308A and the albedo code 308B are given as inputs to the corresponding shape generator 310A and albedo generator 310B, respectively, to generate the shape representation 312A in the shape physical space and the albedo representation 312B in the albedo physical space. In some exemplary embodiments, the shape generator 310A corresponds to a convolutional generator that generates the 3D shape 312A of image 304 from the shape latent space 308A. For example, the generated 3D shape 312A may consist of three channels in a texture mapping space, such as UV space, representing the 3D coordinates of the vertices of the 3D shape 312A. In some exemplary embodiments, the albedo representation 312B is represented as an albedo map, which can be used to determine the albedo at a specific or arbitrary location on the 3D shape 312A. For example, the albedo representation 312B may correspond to a Red-Green-Blue (RGB) albedo map in UV space. In some embodiments, the albedo representation 312B includes one or more channels for specular reflection. For example, the albedo representation 312B may correspond to a red-green-blue specular albedo map having four channels (red diffuse, green diffuse, blue diffuse, and specular) in UV space. Different variations of the albedo code 308B in the albedo latent space may correspond to different variations of the albedo map of the albedo representation 312B in the albedo physical space for achieving a photorealistic image in the final output image of the image manipulation. To that end, in some embodiments, the albedo generator 310B corresponds to a style-based generative adversarial network (GAN).
[0041] In some exemplary embodiments, the illumination physical space 308D may be represented using spherical harmonic parameterization. Spherical harmonic parameterization allows for the reconstruction of physical properties corresponding to illumination (e.g., illumination 104D) illuminating a 3D object in image 102, which can approximate arbitrary unknown illumination conditions. In some exemplary embodiments, the pose physical space 308C may correspond to a 6-degree-of-freedom (6-DOF) pose vector, including three parameters for 3D rotation via axial angle representation and three parameters for 3D translation.
[0042] The generated 3D shape 312A, albedo map 312B, pose physical space 308C, and illumination physical space 308D are given as input to the differentiable renderer 314. The differentiable renderer 314 uses the physical space pose representation 308C and the physical space illumination representation 308D to generate a rendered image 316 of the 3D face model from the 3D shape 312A and albedo representation 312B.
[0043] The process, starting with processing image 102, generating a masked image 304, extracting one or more corresponding physical properties 104A to 104D from the pixels using encoders 306A to 306D, generating shape representations 312A and albedo representations 312B, and outputting a rendered image 316, corresponds to the face reconstruction pipeline 350 shown in Figure 3.
[0044] Furthermore, the rendered image 316 of the 3D face model is masked using the mask image 302 to produce a reconstructed, masked and rendered face image 318. The reconstructed, masked and rendered face image 318 is combined with a reconstructed, 2D masked hair image 330. In parallel, the masked hair image 330 is reconstructed from image 102 using a hair manipulation algorithm. In some embodiments, the hair manipulation algorithm corresponds to an encoder / generator architecture including a hair encoder 320 and a hair generator 324. The encoder 320 extracts hair features 322 corresponding to hair from image 102. The extracted hair features 322 are provided as input to the hair generator 324. The hair generator 324 generates a 2D hair image, such as hair image 326. Hair image 326 is masked with a hair mask 328 to produce a masked hair image 330.
[0045] The combination of the masked and rendered face image 318 and the masked hair image 330 generates a combined composite face image 332 from the face in image 102. In some embodiments, the reconstructed combined composite face image 332 is input to a refiner network 334 to produce a photorealistic refined composite face image 336 as the output of the image manipulation of image 102.
[0046]
number
[0047] In some exemplary embodiments, the albedo generator 310B corresponds to a style-based generative adversarial network (style GAN) generator. To that end, the system 200 can be adversarially trained using a GAN framework with a set of photographic images of human faces to generate a photorealistic output image 336 that exhibits different variations of the human face in image 102.
[0048] Style-based GAN generators can be used to generate realistic images of objects from a predetermined class. These realistic images can be generated by allowing the input to modulate the intermediate layers of a generator neural network (e.g., an albedo generator 310B). The output space of a style-based GAN generator can be the same as the space of real-world training data, i.e., the space of 2D images of objects such as faces. Therefore, a style-based GAN applied to an adversarial discriminator network in the space of 2D images of objects can be trained to directly generate photorealistic 2D images.
[0049] In some embodiments, one or more discriminator networks may be used to train the style-based GAN architecture of the albedo generator 310B, as well as other generators in the framework 300, such as one or more of the shape generator 310A and the hair generator 324. The discriminator network may be applied at one or more stages of the framework 300. In some embodiments, for example, the discriminator network is applied to the output or intermediate image of the framework 300 after the differentiable renderer, such as one or more of the rendered image 316, the masked and rendered face image 318, the combined synthetic face image 332, and the refined synthetic face image 336.
[0050] In some embodiments, part or all of the neural network 210, for example, the face reconstruction pipeline 350, may be pre-trained. This is further illustrated in Figure 4.
[0051] Figure 4 shows a flowchart illustrating a process 400 for manipulating a 2D image of a predetermined class of 3D objects, according to an exemplary embodiment of the present disclosure. The process 400 begins with step 402. The steps of process 400 may be performed by processor 206.
[0052] Step 404 corresponds to pre-training the face reconstruction pipeline 350 of the framework 300 in Figure 3. The face reconstruction pipeline 350 is performed by encoders and generators, such as encoders 306A-306D and shape generator 310A and albedo generator 310B of the framework 300. For that purpose, in step 404, the encoders 306A-306D and generators 310A and 310B of the neural network 210 are pre-trained using a pre-training dataset. For example, the face reconstruction pipeline 350 is trained to take an image from the pre-training dataset as input image 304 and output a rendered image, such as rendered image 316, as a reconstruction of the input image 304. In some embodiments, the pre-trained input image 304 and pre-trained output image 316 are not masked. The face reconstruction pipeline 350 may be pre-trained using a dataset of synthetically generated face images, such as face images generated by a linear 3D morphable model (3DMM). In such pre-training, the face reconstruction pipeline 350 learns to capture facial features using synthetically generated face images. For example, 80,000 2D face images for pre-training may be rendered by randomly sampling a 3D head model from a 3DMM, such as a Faces Learned with an Articulated Model and Expressions (FLAME) model, and rendering the sampled model into a 2D image. Before each sampled 3D model is rendered to generate training set images, it may be converted to 3D so that the resulting 2D face images in the training set correspond to the same 2D alignment as the faces in the actual face image dataset. The actual face image dataset is also used during training step 406. Furthermore, ground truth values for one or more isolated physical properties, such as ground truth values corresponding to one or more physical properties 104A to 104D, are used to generate each of the images in the pre-training set.
[0053]
number
[0054]
number
[0055] During training 406, the neural network 210 is trained on real 2D images, such as image 102, to generalize the model pre-trained on synthetic face images in step 404 to work with real face images. In some exemplary embodiments, the neural network 210 may be trained using a dataset of face photographs and images. In some cases, the face photographs and images may include faces with accessories such as glasses. In such cases, accessories such as glasses may be excluded from the dataset. Face mask image 302 (M f ) are automatically obtained for each image in the dataset using a semantic segmentation network. In addition to or instead of this, face mask images 302 may be obtained using different methods such as face analysis, facial landmark localization, or manual annotation by a human. The masked 2D face images 304 are input to the neural network 210. In some embodiments, during the training process 406, the ground truth values of one or more physical characteristics 104A-104D for the actual face images are unknown, and the reconstruction loss is not applied to one or more physical characteristics 104A-104D.
[0056] During training 406, the loss function may include a pixel-by-pixel reconstruction loss between the masked input face image 304 and the reconstructed face image 318, an identity loss that measures the identity similarity between the masked input face image 304 and the reconstructed masked and rendered face image 318, a marker loss that measures the distance between the image plane projection of the 3D face marker positions in the input image 102 and the corresponding positions in the 3D shape model 312A, and a non-saturated logistic GAN loss to improve realism in the masked and rendered image 318.
[0057] The pixel-by-pixel reconstruction loss between the masked input face image 304 and the reconstructed face image 318 is given by equation (8).
[0058]
number
[0059]
number
[0060]
number
[0061] In some cases, during training of the neural network 210 for the extraction of one or more physical properties 104A-104D, there may be ambiguity in the extracted physical properties 104A-104D. This ambiguity may correspond to the relative contribution of color illumination intensity and surface albedo to the RGB appearance of skin pixels in an image such as image 102. For this purpose, the ambiguity can be overcome based on the albedo normalization loss corresponding to the albedo map 312B.
[0062]
number
[0063] The albedo normalization loss (Equation 11) can minimize the projection of the albedo map 312B from a 3DMM (e.g., a FLAME model) trained on actual face albedo scans to the primary constituent space for albedo. In addition to, or instead of, to address the same ambiguity,
[0064]
number
[0065] The illumination normalization loss, which can be expressed as follows, may be used.
[0066]
number
[0067] Furthermore, the shape code 308A(α) and the albedo code 308B(β) are normalized. The normalization of the shape code 308A(α) and the albedo code 308B(β) is as follows:
[0068]
number
[0069] It is expressed as follows. In step 408, the masked and rendered face image 318 is combined with the masked hair image 330 to obtain a combined composite face image 332. In some embodiments, the hair manipulation pipeline modifies the appearance of the hair to correspond to the modified pose of the face, which is further illustrated in Figure 5.
[0070] During refinement 410, the combined synthetic face image 332 is further processed by the refiner neural network 334 to produce a refined synthetic face image 336. The combined synthetic face image 332 may not exhibit sufficient variation in areas such as the eye region and may lack certain details such as eyelashes, facial hair, and teeth. Furthermore, because the face and hair are processed separately, some combined synthetic face images, such as synthetic face 332, may have a problem of merging the face and hair. For this purpose, the refiner network 334 generates a photorealistic refined synthetic face image 336 that overcomes the gap in realism between the combined synthetic face images 332 and their corresponding original images 102, with minimal modifications made to the reconstruction. The refined synthetic image 336 is more photorealistic than the combined synthetic face image 332, but the difference from the combined synthetic face image 332 is minimal. In some exemplary embodiments, the refiner neural network 334 may correspond to a convolutional neural network architecture such as a U-net. The U-net manipulates the reconstructed combined synthetic image 332 and outputs the final image 336. The refined neural network 334 can be trained using pairs of the original image from the training dataset and the reconstructed face image 332.
[0071] The refined neural network 334 can be trained using a set of loss functions. These loss functions may include a perceptual loss that measures the perceptual similarity between the refined synthetic image 336 and the input image 102, an identity loss that measures the identity similarity between the refined synthetic image 336 and the input image 102, and a non-saturated logistic GAN loss to improve the photorealism of the refined synthetic image 336.
[0072] In step 412, the reconstructed, refined composite image 336 is produced as an output containing a realistic portrait of a human face based on image 102. In some alternative embodiments, the refined composite image 336 may be produced as an anonymized version of the face image in the input image 102. The anonymized version may be produced using an anonymization technique such as a similarity-based metric. The produced anonymized version may allow a photograph or video of one or more people in an input image, such as the input image 102, to be shared to a public forum where the user may restrict the sharing of actual face images. In such a case, the user may share an anonymized version corresponding to the refined composite image 336 on the public forum while protecting the user's actual identity. In step 414, process 400 ends.
[0073] Figure 5 shows a framework corresponding to a hair manipulation pipeline 500 according to an exemplary embodiment of the present disclosure. In some exemplary embodiments, the hair manipulation pipeline 500 corresponds to manipulating one or more of the hair image 326, the hair mask 328, and the masked hair image 330. The hair manipulation algorithm 500 may include a Multi-Input-Conditioned Hair Image GAN (MichiGAN) that processes hair attributes such as hair shape, structure, and appearance from image 102 in a separated manner. Hair shape corresponds to a 2D binary mask of the hair region in image 102. Structure represents a 2D hair orientation map. Appearance refers to the overall color and style of the hair, which may be encoded as a hair appearance code 504 in the hair appearance latent space.
[0074] In some exemplary embodiments, the hair manipulation pipeline 500 may be used to manipulate the shape and structure of hair from an input image 102 without changing the color and style of the hair from the original image 102. To that end, the hair manipulation algorithm 500 extracts hair codes 504 in the hair appearance latent space from the input image 102.
[0075] In some exemplary embodiments, the hair manipulation pipeline 500 is coupled with the face reconstruction pipeline 350. The face reconstruction pipeline 350 is used to extract a 3D model 502 of the face in the image 102. The extracted 3D model 502 includes information about one or more physical properties 104A-104D, such as one or more of the shape code 308A, albedo code 308B, shape representation 312A, albedo representation 312B, pose representation 308C, and lighting representation 308D.
[0076] In some exemplary embodiments, the hair manipulation pipeline 500 is performed iteratively, for example, to achieve larger pose operations by manipulating the pose through several intermediate steps, each of which realizes a smaller pose operation. The iterative approach for the hair manipulation pipeline 500 can improve the photorealistic output of the hair image 326, particularly for larger pose operations. Once the 3D model 502 and hair codes 504 are extracted from the input image 102, they remain constant throughout all iterations of the hair manipulation pipeline 500.
[0077] In each iteration, the hair manipulation algorithm 500 may use a reference pose 504A from the previous iteration and a target pose 506A. In the first iteration, the reference pose 504A may be pose 308C extracted from the input image 102. In the last iteration, the target pose 506A may be the final desired pose for the final output image. In each intermediate iteration (referred to as iteration t), the reference pose 504A is equal to the target pose from the previous iteration (i.e., iteration t-1), and the target pose 506A is one step closer to the final desired pose for the final output image. In each iteration, a 2D strain field 508 is calculated using the reference pose 504A and the target pose 506A. The 2D strain field 508 represents how the image pixels corresponding to the hair move in 2D as a result of the 3D head model 502 transitioning from reference pose 504A to target pose 506A. The 2D distortion field 508 can be calculated based on the projection of the 3D vertices of the 3D model 502 onto a 2D image plane (for example, the 2D image plane of the input image 102) as the pose of the 3D model 502 changes from a reference pose 504A to a target pose 506A. For example, as shown in Figure 5, the 3D vertices of the 3D model 502 in the reference pose 504A may be represented by the face mesh 504B, and the 3D vertices of the 3D model 502 in the target pose 506A may be represented by the face mesh 506B. The 2D distortion field 508 can be obtained by extrapolating the distortion field at image locations that are part of the face to image locations that are not part of the face.
[0078] The reference image 510 input to each iteration is a face image or rendering in reference pose 504A, while the target image 520 output from each iteration is a rendering of the face in target pose 506A. In the first iteration, the reference image 510 may be the original input image 102. In each intermediate iteration, iteration t, the output image 520 from the previous iteration t-1 may be used as the input reference image 510 for iteration t. Alternatively, or in addition to this, the combined image 518 from the previous iteration t-1 may be used as the reference image 510 for iteration t. The distortion field 508 is used to distort the reference mask 512A and reference orientation map 512B of the reference image 510. The reference image 510, along with the reference mask 512A and reference orientation 512B, may be obtained from the previous iteration of the hair manipulation algorithm 500. The distortion field 508 distorts the reference mask 512A to create a distorted mask 514A, and distorts the reference orientation map 512B to create a distorted orientation map 514B. The distorted mask 514A and the distorted orientation map 514B are normalized to obtain the corresponding target mask 516A and target orientation 516B, respectively. The hair manipulation algorithm 500 combines the target mask 516A and target orientation 516B with the hair appearance latent space 504 to generate a target image 518. The target image 518 is input to the refiner neural network 334 to generate a photorealistic refined target image 520. The refined target image 520 from iteration t of the hair manipulation algorithm can be used as the reference image 510 in the next iteration, iteration t+1. In the final iteration, the refined target image may correspond to the reconstructed 2D image 336.
[0079] Each of the following is updated in every iteration: base pose 504A, hair appearance code 504, face mesh 504B, face mesh 506B, 2D distortion field 508, base image 510, base mask 512A, base orientation map 512B, distorted mask 514A, base orientation map 512B, distorted orientation map 514B, target mask 516A and target orientation 516B, target image 518, and refined target image 520.
[0080] Figures 6A and 6B together illustrate a method 600 for manipulating 2D images of a predetermined class of 3D objects, according to an exemplary embodiment of the present disclosure. Method 600 begins with operation 602 and ends with operation 606, which is performed by system 200.
[0081] In operation 602, method 600 includes receiving a 2D input image (e.g., image 102) of a predetermined class of 3D input objects and one or more operation commands for manipulating one or more physical properties (e.g., one or more physical properties 104A to 104D) of the 3D input object. One or more physical properties include the 3D shape of the 3D input object, the albedo of the 3D input object, the pose of the 3D input object, and the illumination illuminating the 3D input object. In some exemplary embodiments, the object is selected from a predetermined class of human faces.
[0082] In operation 604, method 600 includes presenting a 2D input image to a neural network (e.g., neural network 210) while manipulating the latent space, physical space, or both of one or more physical properties of a 3D input object according to one or more operation commands. In some embodiments, the neural network is trained to reconstruct the 2D input image by using multiple subnetworks to separate the physical properties of the 3D input object from the pixels of the 2D input image, and by combining the separated physical properties generated by the multiple subnetworks using a differentiable renderer to form a 2D output image. Each of the multiple subnetworks separates one or more of the one or more physical properties of the 3D object. In some embodiments, there is one subnetwork for each physical property of the 3D input object, and each subnetwork of the multiple subnetworks is trained to extract the corresponding physical property from the pixels of the 2D input image into one or both of the latent space and / or physical space of the corresponding physical property. An albedo subnetwork, trained to extract the albedo of a 3D input object from pixels of a 2D input image, has a style-based generative adversarial network (GAN) architecture trained using a set of photographic images of objects from a predetermined class. Operations are performed on isolated physical properties generated by multiple subnetworks to modify the 3D input object before a differentiable renderer reconstructs the 2D output image. In some exemplary embodiments, multiple realistic composite images can be generated from a 2D image of a 3D object. These multiple realistic composite images may include variations in one or more physical properties of the 3D object, such as pose, facial expression, and lighting.
[0083] In operation 606, method 600 includes outputting a 2D output image. In some exemplary embodiments, the 2D output image may be used in at least one of the following: 3D animation of a 3D object from a 2D input image; hairstyle trial on a human face from a 2D input image; reconstruction of a photorealistic composite image of a human face from a 2D input image; and robot assistance using the reconstructed 2D output image.
[0084] Figure 7 shows a block diagram of a system 700 for manipulating 2D images of a predetermined class of 3D objects, according to an exemplary embodiment of the present disclosure. System 700 corresponds to system 200 in Figure 2. System 700 includes a processor 702 and memory 704. Memory 704 may include random access memory (RAM), read-only memory (ROM), flash memory, or any other suitable memory system. Memory 704 is configured to store a neural network 706. Neural network 706 corresponds to neural network 210. The processor 702 may be a single-core processor, a multi-core processor, a computing cluster, or any number of other configurations.
[0085] System 700 also includes an input interface 720 configured to receive 2D input images of 3D input objects of a predetermined class and one or more operation commands for manipulating one or more physical properties of the 3D input objects. The one or more physical properties correspond to one or more physical properties 104A-104D, including the 3D shape of the 3D input object, the albedo of the 3D shape of the 3D input object, the pose of the 3D input object, and the illumination illuminating the 3D input object. In addition to or instead of this, the input interface 720 may receive 2D images from a camera, such as camera 724. Some examples of camera 724 may include an RGBD camera.
[0086] The processor 702 is configured to present a 2D image to the neural network 706 stored in memory 704. The neural network 706 corresponds to the neural network 210.
[0087] In one implementation, a human-machine interface (HMI) 714 within system 700 connects system 700 to camera 724. In addition to, or instead of, a network interface controller (NIC) 716 may be adapted to connect system 700 to network 718 via bus 712. In one implementation, 2D images may be accessed from image dataset 710 via network 718. Image dataset 710 may be used for pre-training and training of neural network 706.
[0088] In addition, or instead, the system 700 may include an output device 726 configured to output a final image, such as a reconstructed 2D image 336. The output device 726 may be connected to the system 700 via an output interface 722.
[0089] In addition to or instead of the above, the system 700 may include a storage 708 configured to store one or more physical properties of the image, such as one or more physical properties 104A to 104D of the image, a reconstructed image of the image (for example, a reconstructed image after one or more physical properties 104A to 104D have been manipulated), a latent space of one or more physical properties 104A to 104D, and so on.
[0090] The data stored in storage 708 can be accessed via network 718 for further processing. For example, processor 702 may access storage 708 via network 718 to utilize the 2D output image in at least one of the following: 3D animation of a 3D object from a 2D input image, hairstyle trial on a human face from a 2D input image, reconstruction of a photorealistic composite image of a human face from a 2D input image, and robot assistance using the reconstructed 2D output image.
[0091] Figure 8 shows a use case scenario 800 for manipulating 2D images of a predetermined class of 3D objects, according to an exemplary embodiment of the present disclosure. Use case scenario 800 corresponds to an application for generating cartoon or animated versions of a user, such as user 802, using system 700. For generating such animated versions of user 802, system 700 may be trained using an animated training dataset. In some cases, the animated version of user 802 may correspond to an anonymized version of user 802. The application may be installed on user device 804. In the exemplary example scenario, user 802 may capture a 2D image including a front view of user 802's face. The 2D image may be shared to system 700 via a network, such as network 718. System 700 processes the 2D image of user 802. After processing, system 700 generates one or more animated versions 806. One or more animated versions of 806 may include facial images of user 802 in different poses and / or expressions, different animated photos of user 802, etc.
[0092] Figure 9 shows a use case scenario 900 for manipulating a 2D image of a predetermined class of 3D objects, according to another exemplary embodiment of the present disclosure. In use case scenario 900, a 2D image of a user, such as user 902, may be used to realistically try out the appearance of different hairstyles. In the exemplary scenario, a 2D facial image of user 902 is shared with an electronic device 904. In some cases, the electronic device 904 may be remotely connected to a system 700. In some other cases, the system 700 may be embedded within the electronic device 904. The electronic device 904 transmits the 2D image of user 902 to the system 700. The system 700 generates different images of user 902 with corresponding different hairstyles in a photorealistic manner such that the different images remain realistic even if user 902 changes the head pose and the corresponding different hairstyles. User 902 may select a desired hairstyle from the generated different images with different hairstyles.
[0093] Figure 10 shows a use case scenario 1000 for manipulating a 2D image of a predetermined class of 3D objects, according to yet another exemplary embodiment of the present disclosure. Use case scenario 1000 corresponds to the reconstruction of a photorealistic composite image of a human face from a 2D input image. In one exemplary scenario, a surveillance camera may capture a 2D image of a person. The captured 2D image of a person may be a side view of the person's face. In such a scenario, the 2D image may be shared with a system 700. The system 700 may generate different images 1002 of the person from different angles, different poses and / or facial expressions. The different images 1002 may include an image having a reconstructed front view of the person. It may be used by a security guard to track and apprehend a suspect. For example, from the different images 1002, an image 1004 having a front view of the person may be selected.
[0094] Figure 11 shows a use case scenario 1100 for manipulating 2D images of a predetermined class of 3D objects, according to yet another exemplary embodiment of the present disclosure. Use case scenario 1100 corresponds to a robotic system that uses the reconstructed 2D output image. Such a robotic system may be used in an industrial automation application for picking up an object from a conveyor belt 1104. There may be times when lighting conditions for the object are poor, which may hinder the robotic system's ability to pick up the object from the conveyor belt 1104. In such poor lighting conditions, system 700 may generate a reconstructed image of an object, such as object 1102, under poor lighting conditions, to help the robotic system pick up object 1102 from the conveyor belt 1104. For that purpose, a 2D image of a 3D object, such as object 1102, may be captured using device 1106. Device 1106 may include a camera. The 2D image of object 1102 is shared from device 1106 to system 700 for further processing. System 700 may reconstruct a 2D object using manipulated illumination to assist the robotic system in picking up the object 1102 from the conveyor belt 1104.
[0095] Thus, system 700 can be used to manipulate 2D images about 3D objects using a neural network trained in an unsupervised manner. Unsupervised training of the neural network eliminates the need for 3D data, which may be difficult to obtain and / or expensive, improving overall processing and storage efficiency. Furthermore, such manipulation of 2D images helps to realize photorealistic images in an efficient and feasible manner.
[0096] The above description provides only exemplary embodiments and is not intended to limit the scope, availability, or configuration of this disclosure. Rather, the above description of exemplary embodiments will provide a practicable description for realizing one or more exemplary embodiments for those skilled in the art. Various modifications may be made in the function and arrangement of the elements without departing from the spirit and scope of the subject matter disclosed as stated in the appended claims.
[0097] Certain details are given in the above description to provide a complete understanding of the embodiments. However, it will be understood by those skilled in the art that the embodiments can be practiced without these specific details. For example, to avoid obscuring the embodiments by describing them in unnecessary detail, systems, processes, and other elements in the disclosed subject matter may be shown as components in the form of block diagrams. In other cases, to avoid obscuring the embodiments, well-known processes, structures, and techniques may be shown without unnecessary detail. Also, the same reference numerals and names in different drawings refer to the same elements.
[0098] Furthermore, individual embodiments may be described as processes depicted as flowcharts, flow diagrams, data flow diagrams, structural diagrams, or block diagrams. While flowcharts can describe operations as sequential processes, many operations may occur in parallel or simultaneously. In addition, the order of operations may be rearranged. A process may terminate when its operations are completed, but it may have additional steps not described or included in the drawings. Moreover, not all operations in any particular process described may occur in all embodiments. A process may correspond to a method, function, procedure, subroutine, subprogram, etc. If a process corresponds to a function, the termination of that function may correspond to the function returning to a called function or main function.
[0099] Furthermore, embodiments of the disclosed subject matter may be implemented at least partially manually or automatically. Manual or automatic implementations may be performed, or at least assisted, through the use of a machine, hardware, software, firmware, middleware, microcode, a hardware description language, or any combination thereof. If implemented in software, firmware, middleware, or microcode, the program code or code segments for performing the required tasks may be stored in a machine-readable medium. A processor may perform the required tasks.
[0100] The various methods or processes outlined herein may be encoded as software executable on one or more processors employing any one of various operating systems or platforms. In addition, such software may be written using any of several preferred programming languages and / or programming tools or scripting tools, and may be compiled as executable machine language code or intermediate code that runs on a framework or virtual machine. Typically, the functionality of program modules may be combined or distributed as desired in various embodiments.
[0101] Embodiments of the present disclosure may be embodied as the example provided. The operations performed as part of the method may be ordered in any preferred manner. Thus, even though they are shown as sequential operations in the exemplary embodiments, embodiments may be constructed in which the operations are performed in a different order than shown. This may include performing several operations simultaneously. Furthermore, the use of ordinal terms such as “first,” “second,” etc., in a claim to modify a claim element does not in itself imply any priority, superiority, or order of one claim element over another, or the temporal order in which the operations of the method are performed, but is merely used as a label to distinguish these claim elements, in order to distinguish one claim element having a certain name from another element having the same name (except for the use of ordinal terms).
[0102] While this disclosure has been described with reference to a preferred embodiment, it should be understood that various other adaptations and modifications are possible within the spirit and scope of this disclosure. Therefore, the aspect of the appended claims is to cover all such variations and modifications so as to fall within the true spirit and scope of this disclosure.
Claims
1. An image processing system for manipulating a two-dimensional (2D) image of a three-dimensional (3D) object of a predetermined class, the image processing system comprising: an input interface configured to receive a 2D input image of a 3D input object of the predetermined class and one or more manipulation instructions for manipulating one or more physical characteristics of the 3D input object, the one or more physical characteristics including a 3D shape of the 3D input object, an albedo of the 3D input object, a pose of the 3D input object, and illumination irradiating the 3D input object, the image processing system further comprising: a memory configured to store a neural network, the neural network being trained to reconstruct the 2D input image by separating the one or more physical characteristics of the 3D input object from the pixels of the 2D input image using a plurality of sub-networks and combining the separated one or more physical characteristics generated by the plurality of sub-networks using a differentiable renderer into a 2D output image, each sub-network extracting one or more of the one or more physical characteristics of the 3D input object, each sub-network of the plurality of sub-networks being trained to extract a corresponding physical characteristic from the pixels of the 2D input image into one or a combination of a latent space and a physical space of the corresponding physical characteristic, the albedo sub-network trained to extract the albedo of the 3D input object from the pixels of the 2D input image having a style-based adversarial generation network (GAN) architecture trained using a set of photographic images of objects from the predetermined class, the image processing system further comprising: including a processor, the processor presenting the 2D input image received from the input interface to the neural network and configured to operate on the latent space, the physical space, or both of the one or more physical characteristics of the 3D input object according to the one or more operation instructions, the operation being performed on the one or more separated physical characteristics generated by the plurality of sub-networks to modify the 3D input object before the differentiable renderer reconstructs the 2D output image, the image processing system further An image processing system including an output interface configured to output the 2D output image. **Claim 2** The image processing system according to claim 1, wherein the neural network further includes a refinement neural network trained to improve the visual representation of the modified 3D input object in the 2D output image. **Claim 3** The image processing system according to claim 1, wherein the object is selected from a predefined class of human faces. **Claim 4** The image processing system according to claim 3, wherein the processor is configured to execute a hair operation pipeline for operating on the hair of the human face separately from the operation of the one or more physical characteristics of the human face. **Claim 5** The processor is configured to execute the hair operation pipeline to modify the hair representation in the 2D output image in response to an operation on the pose of the human face in the 2D input image, and the hair operation pipeline modifies the appearance of the hair to correspond to the modified pose of the human face. The image processing system according to claim 4. **Claim 6** The neural network is pre-trained using a data set including a plurality of synthetically generated human face images. The image processing system according to claim 1. **Claim 7** The processor is configured to generate a plurality of realistic synthetic images from the 2D image of the 3D object, the plurality of realistic synthetic images including variations in the one or more physical characteristics of the pose, expression, and illumination of the 3D object. The image processing system according to claim 1. **Claim 8** The image processing system according to claim 1, wherein the processor is configured to generate a mask image for the 2D input image using a semantic segmentation network.
9. The 2D output image is utilized in at least one of: 3D animation of the 3D object from the 2D input image, a hairstyle trial for a human face in the 2D input image, reconstruction of a photorealistic composite image of the human face from the 2D input image, and a robot system using the reconstructed 2D output image. The image processing system according to claim 1.
10. The 2D output image is used as at least one of: an anonymized version of the 2D input image, and a training data sample image for a different system. The image processing system according to claim 1.
11. A method for manipulating a two-dimensional (2D) image of a three-dimensional (3D) object of a predetermined class, the method comprising: receiving a 2D input image of a 3D input object of the predetermined class and one or more operation instructions for manipulating one or more physical characteristics of the 3D input object, the one or more physical characteristics including the 3D shape of the 3D input object, the albedo of the 3D shape of the 3D input object, the pose of the 3D input object, and the illumination irradiating the 3D input object, the method further comprising: presenting the 2D input image to a neural network and manipulating the latent space, the physical space, or both of the one or more physical characteristics of the 3D input object according to the one or more operation instructions, the manipulation being performed on the separated one or more physical characteristics generated by the plurality of sub-networks to modify the 3D input object before a differentiable renderer reconstructs the 2D output image. The neural network separates one or more of the physical characteristics of the 3D input object from the pixels of the 2D input image using the plurality of sub-networks, and combines the separated physical characteristics generated by the plurality of sub-networks using the differentiable renderer into a 2D output image, thereby reconstructing the 2D input image. Each sub-network extracts one or more of the physical characteristics of the 3D input object, and each sub-network of the plurality of sub-networks is trained to extract a corresponding physical characteristic from the pixels of the 2D input image into one or a combination of the latent space and the physical space of the corresponding physical characteristic. The albedo sub-network trained to extract the albedo of the 3D shape of the 3D input object from the pixels of the 2D input image has a style-based adversarial generation network (GAN) architecture trained using a set of photographic images of objects from the predetermined class, and the method further includes A method including the step of outputting the 2D output image. Claim 12 Using the refiner neural network of the neural network to improve the visual representation of the modified 3D input object in the 2D output The method according to claim 11, further comprising the step of improving the visual representation of the modified 3D input object in the force image. Claim 13 The method according to claim 11, wherein the object corresponds to a human face of a predetermined class. Claim 14 The method according to claim 13, further comprising the step of executing a hair manipulation pipeline for manipulating the hair of the human face, separately from the manipulation of the physical characteristics of the human face. Claim 15 The method according to claim 14, further comprising the step of executing the hair manipulation algorithm to modify the representation of the hair in the 2D output image corresponding to the manipulation of the pose of the input human face in the 2D input image, and the hair manipulation algorithm is trained to synchronize the appearance of the hair with the modified pose of the input human face. Claim 16 The method according to claim 11, further comprising the step of pre-training the neural network using a dataset including synthetically generated human face images. Claim 17 The method according to claim 11, further comprising the step of generating a plurality of realistic composite images from the 2D image of the 3D object, wherein the plurality of realistic composite images include variations in the one or more physical characteristics of the pose, expression, and illumination of the 3D object.
18. The method according to claim 11, further comprising the step of generating a mask image for the 2D input image using a semantic segmentation network.
19. The method according to claim 11, further comprising the step of utilizing the 2D output image in at least one of a 3D animation of the 3D object from the 2D input image, a hairstyle trial for a human face in the 2D input image, a reconstruction of a photo-realistic composite image of the human face from the 2D input image, and a robotic system using the reconstructed 2D output image.
20. The method according to claim 11, further comprising the step of utilizing the 2D output image as at least one of an anonymized version of the input image and a training data sample image for a different method.