Method and system for generating latent texture proxies for object category modeling
Through the 3D proxy geometric structure and neural texture model, the problem of inaccurate rendering of transparent and reflected objects on 3D displays is solved, real-life display effects and real-time updates are achieved.
Patent Information
- Application Number
- CN202080007948.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-30
- Filing Date
- 2020-08-04
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2040-08-04
AI Technical Summary
The prior art is difficult to accurately render objects with transparent, reflective characteristics or complex geometric structures on 3D displays, such as glasses, jewelry, etc., resulting in unrealistic display effects.
Using 3D proxy geometry and neural texture model, a realistic 3D image is generated by generating multiple three-dimensional proxy geometry and neural textures of the object, combined with neural renderers and α masks.
It realizes accurate rendering of transparent and reflected objects on a 3D display, ensuring the realistic display effect and real-time updates that adapt to user movement.
Smart Images

Figure CN114175097B_ABST
Abstract
Description
Technical Field
[0001] This application claims the benefit of U.S. Provisional Application No. 62 / 705,500, filed on June 30, 2020, entitled “GENERATIVE LATENT TEXTURED PROXIES FOR OBJECT CATEGORY MODELING,” the disclosure of which is incorporated herein by reference in its entirety. Technical Field
[0003] This description generally relates to methods, devices, and algorithms for generating content for presentation on a display. Background Art
[0004] A generative model is a machine learning model used to generate data consistent with training data. A generative model can learn a model of a dataset in order to generate data similar to the training data included in the dataset. For example, a generative model can be trained to determine the probability distribution p(X, Y) for a feature X and a label Y in a dataset. Label Y can be provided to a computer system programmed to execute the generative model. In response, the computer system can generate a feature or feature set X that is consistent with label Y. Summary of the Invention
[0005] A system of one or more computers may be configured to perform specific operations or actions by installing software, firmware, hardware, or a combination thereof on the system, which, when operated, causes the system to perform the actions. One or more computer programs may be configured to perform specific operations or actions by including instructions that, when executed by a data processing device, cause the device to perform the actions.
[0006] In one general aspect, systems and methods are described for performing operations utilizing at least one processing device, the operations comprising: receiving a pose associated with an object in image content; generating a plurality of three-dimensional (3D) proxy geometries for the object; generating a plurality of neural textures for the object based on the plurality of 3D proxy geometries, wherein the neural textures define a plurality of different shapes and appearances representing the object; providing the plurality of neural textures to a neural renderer, wherein the plurality of neural textures are provided in a stacked formation; receiving, from the neural renderer and based on the plurality of neural textures, a color image and an alpha mask representing an opacity of at least a portion of the object; and generating a composite image based on the pose, the color image, and the alpha mask.
[0007] These and other aspects may include, alone or in combination, one or more of the following. For example, the method may further include rendering a latent texture onto a target viewpoint based at least in part on a pose associated with the object, wherein each of the plurality of 3D proxy geometries includes a coarse geometric approximation of at least a portion of the object and a latent texture of the object mapped to the coarse geometric approximation. In some embodiments, the plurality of neural textures are configured to reconstruct hidden portions of an object captured in the image content, wherein reconstructing the hidden portions based on a stacked formation of the neural textures enables the neural renderer to generate a transparent layer of the object and a surface behind the transparent layer of the object.
[0008] In some embodiments, each of the plurality of 3D proxy geometries encodes a surface light field associated with an object in the image content, the surface light field including specular reflections associated with the object. In some embodiments, the plurality of neural textures are based at least in part on the pose, the neural textures being generated by: identifying a category of the object; generating a feature map based on the identified category of the object; providing the feature map to a neural network; and generating a neural texture based on a latent code associated with each instance of the identified category and a view associated with the pose. In some embodiments, at least a portion of the object is a transparent material. In some embodiments, at least a portion of the object is a reflective material.
[0009] In some embodiments, the image content includes telepresence image data, the telepresence image data includes at least a user; and the object includes a pair of glasses. In some embodiments, the neural renderer uses a generative model to reconstruct unseen object instances within the identified category, the reconstruction being based on fewer than four captured views of the object. In some embodiments, the synthetic image is generated using a generative latent optimization (GLO) framework and a perceptual reconstruction loss.
[0010] Implementations of the described techniques may include hardware, methods or processes, or computer software on a computer-accessible medium. The details of one or more implementations are set forth in the accompanying drawings and the following description. Other features will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 is a block diagram illustrating an exemplary 3D content system for displaying content on a display device according to implementations described throughout this disclosure.
[0012] Figure 2 is a block diagram of an exemplary system for modeling rendered content in a display device according to implementations described throughout this disclosure.
[0013] Figure 3is a diagram depicting exemplary planar proxies for object classes with well-defined geometric variations according to embodiments described throughout this disclosure.
[0014] Figure 4 is a block diagram of an exemplary network architecture trained by a generative latent optimization framework according to embodiments described throughout this disclosure.
[0015] Figures 5A-5C Examples of simulating, capturing, and extracting image content according to embodiments described throughout this disclosure are illustrated.
[0016] Figure 6 Illustrated are exemplary images of locations that fit based on the models described herein, according to implementations described throughout this disclosure.
[0017] Figures 7A-7C An exemplary virtual try-on application using the models described herein is illustrated according to implementations described throughout this disclosure.
[0018] Figure 8 is a flow chart illustrating one example of a process for generating a synthetic image based on a 3D proxy geometry model according to embodiments described throughout this disclosure.
[0019] Figure 9 Examples of computer devices and mobile computer devices that may be used with the techniques described herein are shown.
[0020] Like reference symbols in the various drawings indicate like elements. DETAILED DESCRIPTION
[0021] Accurate modeling and representation of 3D objects can be challenging when the objects exhibit features such as transparent surfaces, reflective surfaces, and / or thin structures. The systems and techniques described herein can provide a way to model 3D objects with these features using 3D proxy geometry (e.g., texture proxies) to enable accurate rendering of the 3D objects on a 2D or autostereoscopic display (e.g., a 3D display). In some embodiments, the 3D proxy geometry is based on geometric interpolation that constructs the shape of the object within the image content.
[0022] In general, this document describes examples related to modeling the shape and appearance of classes of objects in order to render accurate images depicting 3D objects. In some embodiments, for example, the models described herein can be used to simulate camera-captured objects in a realistic and 3D manner on a screen of a 3D display used, for example, in a multi-way video conference. In some embodiments, the objects can be synthetically generated objects to provide virtual or augmented content within a 3D generated scene. In some embodiments, the objects can be synthetically modified to create randomness and / or realism for a 2D or 3D scene. For example, the models described herein can be used to generate and display objects composed of complex shapes and appearances, some of which may include transparent properties, reflective properties, complex geometric structures, and / or other structural properties that may typically be difficult to depict in 3D.
[0023] For example, conventional display systems may be unable to accurately render complex objects (e.g., eyeglasses, jewelry, reflective clothing, etc.) onto a user captured for 3D display because transparent and / or reflective materials are difficult to reconstruct and render in 3D. The systems and techniques described herein can generate one or more models of specific physical, lighting, and shading aspects of an object (e.g., eyeglasses, jewelry, reflective clothing, and / or non-user-related objects) in order to depict the object in an accurate 3D representation that provides a realistic depiction of the object on a 3D display. In operation, the systems described herein can perform this modeling in real time as the object is captured for rendering in a 3D display. In some embodiments, the systems described herein can perform this modeling and rendering while the user is moving with and / or near the object (i.e., wearing or interacting with the object) during use of the 3D display. In some embodiments, the systems described herein can perform this modeling for other types of objects, including, but not limited to, vehicle parts, painted surfaces, transparent objects, objects filled with liquids, and the like. These objects can be rendered using the modeling and techniques described herein to appear realistic in 3D.
[0024] In some embodiments, the systems and techniques described herein use approximate geometry to generate a general shape and appearance model representing a class of objects to generate a 3D proxy geometry. As used herein, a 3D proxy geometry (texture proxy) represents a rough geometric approximation of a set of objects and a potential texture of one or more objects mapped to the corresponding object geometry. The rough geometry and the mapped potential texture can be used to generate an image of one or more objects in an object class. For example, the systems and techniques described herein can generate a target image on a display by rendering the potential texture onto a target viewpoint and accessing a neural rendering network (e.g., a differentially delayed rendering neural network) to generate an object for 3D telepresence display. In order to learn this potential texture, the system described herein can learn a low-dimensional latent space of neural textures and a shared delayed neural rendering network. The latent space contains all instances of a class of objects and allows interpolation of instances of objects, which can enable reconstruction of instances of objects from several viewpoints.
[0025] To generate the texture of the proxy, the systems and techniques described herein use class-level appearance and geometry interpolation to learn a joint latent space. For example, if the object is an earring, a specific dataset can be selected that includes material reflectance (e.g., for gold, silver, plastic, resin, etc.), earring shape, etc. The proxy can be independently rasterized with its corresponding neural texture and synthesized using a neural network (e.g., U-Net), generating a realistic image and alpha channel (e.g., mapping, mask, etc.) as output. Using the 3D proxy geometry, complex objects can be reconstructed from a sparse set of views (e.g., fewer than four input images).
[0026] In some embodiments, the systems and techniques described herein can evaluate how to display image content captured by a camera for rendering on a 3D display in response to detecting movement of a user accessing the display. For example, if a user (or the user's head or eyes) moves left or right, the systems and techniques described herein can detect such movement to model specific objects within the image capture to determine how to display the object (e.g., image content, user, etc.) in a manner that provides the user of the 3D display with 3D depth, appropriate disparity, and 3D perception of the object. Furthermore, for example, the systems and techniques described herein can be used to provide the same 3D depth, disparity, and perception of the object to other users viewing the object on other 3D displays.
[0027] Figure 1 1 is a block diagram illustrating an exemplary 3D content system 100 for displaying content in a stereoscopic display device according to embodiments described throughout this disclosure. The 3D content system 100 may be used by multiple users to, for example, conduct video conferencing communications (e.g., telepresence sessions) in 3D. Figure 1The system 100 can be used to capture video and / or images of a user during a video conference and use the systems and techniques described herein to model the shape and appearance of 3D objects (e.g., eyeglasses, jewelry, etc.) in order to render accurate images depicting the 3D objects within the video conference session. The system 100 can benefit from using the models described herein because such models can generate and display objects composed of complex shapes and appearances within a video conference, for example, some of which can include transparent properties, reflective properties, complex geometric structures, and / or other structural properties that are generally difficult to depict in 3D.
[0028] like Figure 1 As shown, a first user 102 and a second user 104 are using a 3D content system 100. For example, users 102 and 104 are participating in a 3D telepresence session using the 3D content system 100. In such an example, the 3D content system 100 can allow each of users 102 and 104 to see a highly realistic and visually consistent representation of the other, thereby facilitating the users to interact in a manner similar to being in physical presence with each other.
[0029] Each user 102, 104 may have a corresponding 3D system. Here, user 102 has 3D system 106 and user 104 has 3D system 108. 3D systems 106, 108 may provide functionality related to 3D content, including but not limited to: capturing images for 3D display, processing and presenting image information, processing and presenting audio information. 3D system 106 and / or 3D system 108 may constitute a collection of sensing devices integrated into a unit. 3D system 106 and / or 3D system 108 may include a reference Figure 2 、 4 and some or all of the components described in 9.
[0030] 3D content system 100 may include one or more 2D or 3D displays. Here, 3D display 110 is provided for 3D system 106 and 3D display 112 is provided for 3D system 108. 3D displays 110, 112 may use any of a variety of types of 3D display technologies to provide an autostereoscopic view of 3D to a respective viewer (here, for example, user 102 or user 104). In some embodiments, 3D displays 110, 112 may be standalone units (e.g., self-supporting or hung on a wall). In some embodiments, 3D displays 110, 112 may include or have access to wearable technology (e.g., a controller, a head-mounted display, etc.). In some embodiments, displays 110, 112 may be 2D displays, such as Figures 7A-7C shown.
[0031] Generally speaking, 3D displays, such as displays 110 and 112, can provide images that approximate the 3D optical properties of physical objects in the real world without the use of a head-mounted display (HMD) device. Generally speaking, the displays described herein include a flat panel display, a lenticular lens (e.g., a microlens array), and / or a parallax barrier to redirect the image to multiple different viewing zones associated with the display.
[0032] In some embodiments, the displays 110, 112 can include high-resolution, glasses-free, lenticular three-dimensional displays. For example, the displays 110, 112 can include a microlens array (not shown) comprising a plurality of lenses (e.g., microlenses), wherein a glass spacer is coupled (e.g., bonded) to the microlenses of the display. The microlenses can be designed such that, from a selected viewing position, a user's left eye of the display can view a first set of pixels, while the user's right eye can view a second set of pixels (e.g., wherein the second set of pixels is mutually exclusive with the first set of pixels).
[0033] In some exemplary 3D displays, there may be a single location that provides a 3D view of the image content (e.g., a user, an object, etc.) provided by such a display. The user may sit in this single location to experience appropriate parallax, minimal distortion, and realistic 3D images. If the user moves to a different physical location (or changes head position or eye gaze position), the image content (e.g., the user, objects worn by the user, and / or other objects) may begin to appear less realistic, 2D, and / or distorted. The systems and techniques described herein can reconfigure the image content projected from the display to ensure that the user can move around but still experience appropriate parallax, low distortion, and realistic 3D images in real time. Thus, the systems and techniques described herein provide for maintaining and providing 3D image content of an object for display to the user regardless of any user movement that occurs while the user is viewing the 3D display.
[0034] like Figure 1 As shown, 3D content system 100 can be connected to one or more networks. Here, network 114 is connected to 3D system 106 and 3D system 108. Network 114 can be a publicly available network (e.g., the Internet) or a private network, to name just two examples. Network 114 can be wired, wireless, or a combination of both. Network 114 can include or utilize one or more other devices or systems, including, but not limited to, one or more servers (not shown).
[0035] 3D systems 106 and 108 may include a number of components related to the capture, processing, transmission or reception of 3D information and / or the presentation of 3D content. 3D systems 106 and 108 may include one or more cameras for capturing image content to be included in the 3D presentation. Here, 3D system 106 includes cameras 116 and 118. For example, camera 116 and / or camera 118 may be substantially disposed within the housing of 3D system 106, such that the objective lens or lens of the respective camera 116 and / or 118 captures image content through one or more apertures in the housing. In some embodiments, cameras 116 and / or 118 may be separate from the housing, such as in the form of standalone devices (e.g., having a wired and / or wireless connection to 3D system 106). Cameras 116 and 118 may be positioned and / or oriented to capture a sufficiently representative view of a user (e.g., user 102). While cameras 116 and 118 generally do not obstruct the view of 3D display 110 for user 102, the layout of cameras 116 and 118 may be arbitrarily selected. For example, one of cameras 116 and 118 may be positioned somewhere above user 102's face, while the other may be positioned somewhere below the face. For example, one of cameras 116 and 118 may be positioned somewhere to the right of user 102's face, while the other may be positioned somewhere to the left. For example, 3D system 108 may include cameras 120 and 122 in a similar manner. Additional cameras are possible. For example, a third camera may be placed near or behind display 110.
[0036] 3D systems 106 and 108 may include one or more depth sensors to capture depth data for use in 3D presentations. Such depth sensors can be considered part of the depth capture component of 3D content system 100, used to characterize the scene captured by 3D systems 106 and / or 108 for accurate representation of the scene on a 3D display. Furthermore, the system can track the position and orientation of the viewer's head so that the 3D presentation can be rendered with an appearance corresponding to the viewer's current viewpoint. Here, 3D system 106 includes depth sensor 124. Similarly, 3D system 108 may include depth sensor 126. Any of a variety of types of depth sensing or depth capture can be used to generate depth data. In some embodiments, assisted stereo depth capture is performed. For example, a light point can be used to illuminate the scene, and stereo matching can be performed between two corresponding cameras. This illumination can be accomplished using waves of a selected wavelength or range of wavelengths. For example, infrared (IR) light can be used. For example, in some embodiments, a depth sensor may not be utilized when generating a view on a 2D device. The depth data may include or be based on any information about a scene that reflects the distance between a depth sensor (e.g., depth sensor 124) and an object in the scene. For content in an image corresponding to an object in the scene, the depth data reflects the distance (or depth) to the object. For example, the spatial relationship between a camera and a depth sensor may be known and may be used to correlate an image from the camera with a signal from the depth sensor to generate depth data for the image.
[0037] Images captured by the 3D content system 100 may be processed and thereafter displayed as a 3D presentation. Figure 1 As depicted in the example of , a 3D image 104 ′ with an object (glasses 104 ″) is presented on a 3D display 110 . Thus, the user 102 may perceive the 3D image 104 ′ and the glasses 104 ″ as a 3D representation of the user 104 , which may be located away from the user 102 . The 3D image 102 ′ is presented on a 3D display 112 . Thus, the user 104 may perceive the 3D image 102 ′ as a 3D representation of the user 102 .
[0038] 3D content system 100 can allow participants (e.g., users 102, 104) to engage in audio communications with each other and / or other people. In some embodiments, 3D system 106 includes a speaker and a microphone (not shown). For example, 3D system 108 can similarly include a speaker and a microphone. In this way, 3D content system 100 can allow users 102 and 104 to participate in a 3D telepresence session with each other and / or other people.
[0039] Figure 2is a block diagram of an exemplary system 200 for modeling content for rendering in a 3D display device according to embodiments described throughout this disclosure. System 200 can be used as or included in one or more embodiments described herein, and / or can be used to perform one or more examples of 3D processing, modeling, or rendering described herein. The entire system 200 and / or one or more of its individual components can be implemented according to one or more examples described herein.
[0040] System 200 includes one or more 3D systems 202. In the depicted example, 3D systems 202A, 202B, through 202N are shown, where index N indicates an arbitrary number. 3D system 202 may provide for the capture of visual and audio information for 3D presentation and forward the 3D information for processing. Such 3D information may include an image of a scene, depth data about the scene, and audio from the scene. For example, 3D system 202 may be used as or included in 3D system 106 and 3D display 110 ( Figure 1 )middle.
[0041] System 200 may include multiple cameras, as indicated by camera 204. Any type of light-sensing technology may be used to capture images, such as the type of image sensors used in common digital cameras. Cameras 204 may be of the same type or of different types. For example, the cameras may be positioned anywhere on a 3D system, such as 3D system 106.
[0042] System 202A includes a depth sensor 206. In some embodiments, depth sensor 206 operates by broadcasting an IR signal onto a scene and detecting a response signal. For example, depth sensor 206 can generate and / or detect light beams 128A-B and / or 130A-B.
[0043] System 202A also includes at least one microphone 208 and speaker 210. For example, these can be integrated into a head-mounted display worn by the user. In some embodiments, microphone 208 and speaker 210 can be part of 3D system 106 rather than part of the head-mounted display.
[0044] The system 202 further includes a 3D display 212 capable of presenting 3D images in a stereoscopic manner. In some embodiments, the 3D display 212 can be a stand-alone display, and in some other embodiments, the 3D display 212 can be included in a head-mounted display unit configured to be worn by a user to experience 3D presentations. In some embodiments, the 3D display 212 operates using parallax barrier technology. For example, a parallax barrier can include parallel vertical stripes of a substantially opaque material (e.g., an opaque film) placed between the screen and the viewer. Due to the parallax between the viewer's respective eyes, different portions of the screen (e.g., different pixels) are viewed by the corresponding left and right eyes. In some embodiments, the 3D display 212 operates using lenticular lenses. For example, alternating rows of lenses are placed in front of the screen that direct light from the screen toward the viewer's left and right eyes, respectively.
[0045] System 200 may include a server 214 that may perform certain tasks of data processing, data modeling, data coordination, and / or data transmission. Server 214 and / or its components may include reference Figure 9 Some or all of the components described.
[0046] Server 214 includes a 3D content generator 216 that may be responsible for rendering 3D information in one or more ways. This may include receiving 3D content (e.g., from 3D system 202A), processing the 3D content, and / or forwarding the (processed) 3D content to another participant (e.g., to another 3D system 202).
[0047] Some aspects of the functionality performed by 3D content generator 216 may be implemented by shader 218. Shader 218 may be responsible for applying shadows to certain portions of an image and may also perform other services related to images that have been or will be provided with shadows. For example, shader 218 may be utilized to offset or hide some artifacts that might otherwise be generated by 3D system 202.
[0048] Shading refers to one or more parameters that define the appearance of image content, including but not limited to the color, surface, and / or polygons of objects in the image. In some embodiments, shading can be applied or adjusted to one or more portions of image content to change how those portions of image content appear to the viewer. For example, shading can be applied / adjusted to make portions of image content darker, lighter, more transparent, etc.
[0049] The 3D content generator 216 can include a depth processing component 220. In some implementations, the depth processing component 220 can apply shading (e.g., darker, lighter, transparent, etc.) to the image content based on one or more depth values associated with the image content and based on one or more received inputs (e.g., content model inputs).
[0050] The 3D content generator 216 can include an angle processing component 222. In some embodiments, the angle processing component 222 can apply shadows to image content based on the orientation (e.g., angle) of the content relative to the camera capturing the image content. For example, shadows can be applied to content that is facing away from the camera at angles above a predetermined threshold. This can allow the angle processing component 222 to reduce brightness and fade content as the surface is turned away from the camera, as just one example.
[0051] The 3D content generator 216 includes a renderer module 224. The renderer module 224 can render content to one or more 3D systems 202. The renderer module 224 can, for example, render an output / composite image that can be displayed in the system 202, for example.
[0052] like Figure 2 As shown, server 214 also includes a 3D content modeler 230, which can be responsible for modeling 3D information in one or more ways. This can include receiving 3D content (e.g., from 3D system 202A), processing the 3D content, and / or forwarding the (processed) 3D content to another participant (e.g., to another 3D system 202). 3D content modeler 230 can utilize architecture 400 to model objects, as described in further detail below.
[0053] Gesture 232 can represent a gesture associated with the captured content (e.g., an object, a scene, etc.). In some embodiments, gesture 232 can be detected and / or otherwise determined by a tracking system (not shown) associated with system 100 and / or 200. Such a tracking system can include sensors, cameras, detectors, and / or markers to track the position of all or part of the user. In some embodiments, the tracking system can track the user's position in the room. In some embodiments, the tracking system can track the position of the user's eyes. In some embodiments, the tracking system can track the position of the user's head.
[0054] In some embodiments, the tracking system can track the position of the user (or the position of the user's eyes or head) relative to the display device 212, for example, in order to display images with appropriate depth and parallax. In some embodiments, the head position associated with the user can be detected and used as a direction for simultaneously projecting images to the user of the display device 212 via, for example, microlenses (not shown).
[0055] Class 234 can represent a classification of a particular object 236. For example, class 234 can be glasses and the objects can be blue glasses, clear glasses, round glasses, etc. Any class and object can be represented by the models described herein. Class 234 can be used as a basis for training a generative model on object 236. In some embodiments, class 234 can represent a dataset that can be used to synthetically render a class of 3D objects from different viewpoints, thereby accessing a set of real poses, color space images, and masks for multiple objects of the same class.
[0056] A three-dimensional (3D) proxy geometry 238 represents a (rough) geometric approximation of a set of objects and a latent texture 239 of one or more objects mapped to the corresponding object geometry. The coarse geometry and the mapped latent texture 239 can be used to generate an image of one or more objects in an object class. For example, the systems and techniques described herein can generate objects for 3D telepresence display by rendering the latent texture 239 onto a target viewpoint and accessing a neural rendering network (e.g., a differentially deferred rendering neural network) to generate a target image on a display. To learn such a latent texture 239, the systems described herein can learn a low-dimensional latent space of neural textures and a shared deferred neural rendering network. The latent space contains all instances of a class of objects and allows interpolation of instances of an object, which enables reconstruction of instances of an object from several viewpoints.
[0057] The neural texture 244 represents the learned feature map 240 that was trained as part of the image capture process. For example, when an object is captured, the neural texture 244 can be generated using the feature map 240 and the 3D proxy geometry 238 for the object. In operation, the system 200 can generate and store the neural texture 244 for a particular object (or scene) as a map on top of the 3D proxy geometry 238 for the object. For example, the neural texture can be generated based on the latent code associated with each instance of the identified class and the view associated with the pose.
[0058] The geometry approximation 246 may represent a shape-based proxy for the geometry of the object. The geometry approximation 246 may be a mesh-based, shape-based (eg, triangle, diamond, square, etc.) free-form version of the object.
[0059] The neural renderer 250 can generate an intermediate representation of an object and / or scene to be rendered, for example, using a neural network. The neural texture 244 can be used in conjunction with a 5-layer U-Net, such as the neural network 242 operating with the neural renderer 250, to jointly learn features about the texture map (e.g., feature map 240). The neural renderer 250 can incorporate view-dependent effects by modeling the difference between the real appearance (e.g., ground truth) and the diffuse reprojection using an object-specific convolutional network. This effect can be difficult to predict based on scene knowledge, and therefore, a GAN-based loss function can be used to render realistic output.
[0060] RGB color channels 252 (e.g., a color image) represent three output channels. For example, the three output channels may include a red channel, a green channel, and a blue channel (e.g., RGB) representing a color image. In some embodiments, color channels 252 may be a YUV mapping that indicates which colors are to be rendered for a particular image. In some embodiments, color channels 252 may be a CIE mapping. In some embodiments, color channels 252 may be an ITP mapping.
[0061] Alpha (α) 254 represents an output channel (e.g., a mask) that represents how a particular pixel color will merge with other pixels when overlapping for any number of pixels in an object. In some embodiments, alpha 254 represents a mask that defines the transparency level (e.g., semi-transparent, opaque, etc.) of the object.
[0062] The above exemplary components are described herein as being implemented in a server 214, which may be accessible via a network 260 (which may be connected to a server 214). Figure 1 114 in the example embodiment) communicates with one or more 3D systems 202. In some embodiments, the 3D content generator 216 and / or its components may alternatively or additionally be implemented in some or all of the 3D systems 202. For example, the modeling and / or processing described above may be performed by a system that originates the 3D information before forwarding it to one or more receiving systems. As another example, the originating system may forward images, modeling data, depth data, and / or corresponding information to one or more receiving systems, which may perform the processing described above. Combinations of these approaches may also be used.
[0063] Thus, system 200 is an example of a system including a camera (e.g., camera 204), a depth sensor (e.g., depth sensor 206), and a 3D content generator (e.g., 3D content generator 216), the system having a processor that executes instructions stored in a memory. Such instructions can cause the processor to use depth data included in the 3D information (e.g., via depth processing component 220) to identify image content in an image of a scene included in the 3D information. The image content can be identified as being associated with depth values that meet a criterion. For example, the processor can generate modified 3D information by applying a model generated by 3D content modeler 230, which can be provided to 3D content generator 216 to appropriately depict the composite image 256.
[0064] Composite image 256 represents a 3D stereoscopic image of a particular object 236 with appropriate parallax and viewing configuration for the eyes associated with a user accessing a display (e.g., display 212) based at least in part on the tracked position of the user's head. For example, each time the user moves their head position while viewing the display, system 200 can be used to determine at least a portion of composite image 256 based on output from 3D content modeler 230. In some implementations, composite image 256 represents object 236 and image content within the view of other objects, the user, or captured object 236.
[0065] In some embodiments, the processor (not shown) of systems 202 and 214 may include a graphics processing unit (GPU) (or communicate with a graphics processing unit (GPU)). In operation, the processor may include or have access to memory, storage devices, and other processors (e.g., CPU). In order to facilitate graphics and image generation, the processor may communicate with the GPU to display images on a display device (e.g., display device 212). The CPU and GPU may be connected via a high-speed bus, such as PCI, AGP, or PCI-Express. The GPU may be connected to a display via another high-speed interface, such as HDMI, DVI, or Display Port. Typically, the GPU may render image content in pixel form. The display device 212 may receive image content from the GPU and may display the image content on a display screen.
[0066] Figure 3is a diagram depicting exemplary planar proxies for classes of objects with well-defined geometric variations, according to embodiments described throughout the present disclosure. For example, planar proxy 302 is depicted as the left side of a pair of glasses 300. Planar proxy 302 represents a planar billboard that models the left side of glasses 300. Similarly, planar proxy 304 is shown as representing the center portion (e.g., the front) of the glasses while planar proxy 306 represents the right side of glasses 300. Glasses 300 represent examples of objects. The systems and techniques described herein can generate and render 3D content using other objects and planar proxy shapes that represent those objects. For example, other proxies can include, but are not limited to, cuboids, cylinders, spheres, triangles, and the like.
[0067] A flat proxy can represent a texture-mapped object (or portion of an object) that can be used as a substitute for complex geometry. Because manipulating and rendering geometry proxies is computationally less complex than manipulating and rendering the corresponding detailed geometry, a flat proxy representation can provide simpler shapes for reconstructing views. A flat proxy representation can be used to generate such views. Using a flat proxy can provide the advantage of low computational cost when attempting to manipulate, reconstruct and / or render objects with highly complex appearances—such as glasses, cars, clouds, trees, and grass, to name a few. Similarly, with the availability of powerful graphics processing units, real-time game engines can provide the use of proxies (e.g., geometry representations) with multiple levels of detail that can be swapped in and out with distance, which use 3D proxy geometry to generate mappings that replace lower-level detail geometry.
[0068] In operation, the system 200 can generate plane proxies 302-304 by computing a bounding box (e.g., a rough visual shell) for each object using the extracted alpha mask. In general, for any number of pixels in the object, the alpha mask represents how a particular pixel color merges with other pixels when superimposed. The system 200 can then specify a region of interest in the image of the glasses. The region of interest can be specified using head coordinates. The system 200 can then extract planes that probabilistically match the surface as observed from the corresponding orthogonal projections. In this example, the planes used to generate the proxies 302-304 are the right view, center view, and left view depicting three sides of the glasses.
[0069] In general, the system 200 can generate planar proxies for any number of images that can be used as training data for a neural network. For example, a neural network can determine how to appropriately display a particular object (e.g., several pairs of glasses) captured by, for example, a camera. Thus, each pair of glasses used as training data for the neural network can be associated with a unique proxy geometry. In some embodiments, during training, the system 200 can detect the pose of an object in an image. In some embodiments, the system 200 can generate views of a particular object by combining a dataset of images with the object and using the detected pose to simulate the object from a viewpoint based on the pose.
[0070] In some embodiments, system 200 can construct a latent space for the glasses and feed the latent space for the glasses to, for example, NN 242, which can then generate a texture map for the glasses. In some embodiments, system 200 can reduce the number of instances of plane proxies from the training data to perform few-shot reconstruction, while using the remaining plane proxies to train the class-level model of the neural network. For example, the remaining plane proxies representing the glasses image can be used to train the glasses category (e.g., category 234) of neural network 242.
[0071] Any number of object categories can be trained for use with the NN 242. For example, the system 200 can train potential 3D proxy geometries using cars, living plants, and / or other categories of objects that may be thin, reflective, transparent, and / or otherwise difficult to accurately model and render in 3D. For example, the system 200 can model a car using free-form 3D proxy geometry and / or geometry meshes based on sampling multiple car objects.
[0072] In another example, a thin object may be captured, such as x-ray film, camera negative, or other film that can be backlit and displayed on a 2D or 3D video. The systems and techniques described herein can employ a planar proxy to appropriately depict and / or correct the image content within the film so that the film (e.g., x-ray, etc.) is properly conveyed to a user viewing the 2D or 3D video.
[0073] Figure 4 is a block diagram of an exemplary network architecture 400 trained by the Generative Latent Optimization framework, according to embodiments described throughout this disclosure. Generally speaking, architecture 400 is an example of utilizing system 200 to parameterize a neural texture using a 3D proxy geometry P using a generative model capable of generating various shapes and appearances of objects. An example using glasses as an exemplary object to be modeled is depicted. However, any object or class of objects may be substituted and used in architecture 400 to model and generate 3D image content.
[0074] As shown in FIG4 , the object set is generated as a map (z) 402, which represents the potential code of each object instance i as z i ∈R n The map (z) 402 of the latent space may be an eight-dimensional (8D) map. The map 402 may include random values optimized using the architecture 400.
[0075] In operation of the architecture 400 (e.g., using the system 200), the map (z) 402 is provided to a multilayer perceptron (MLP) neural network 404 (e.g., NN 242) to generate a plurality of neural textures 244, depicted in this example as neural texture 406, neural texture 408, and neural texture 410. The neural textures 406-410 may represent portions of a mesh that define portions of the geometry and / or texture of a particular object represented in the map (z) 402.
[0076] The MLP NN 404 (e.g., NN 242) can lift elements represented in the 8D map to a higher dimensional space (e.g., 512 dimensions). The architecture 400 utilizes a pose 412 associated with the captured image (e.g., a pose of an agent generated by the captured image) to generate neural textures 406-408, samples 414, 416, and 418 and corresponding depths 420, 422, 424, and corresponding normal viewpoints 426, 428, and 430.
[0077] Assuming a collection of objects of a particular class, the system 200 defines the potential code for each instance i as z i ∈R n The model described herein and used by the architecture 400 can generate and use a set of K agents {P i ,1,...,P i,K} a rough geometry (i.e., a triangle mesh with UV coordinates). For example, the architecture 400 can project a 2D image onto the surface of a 3D proxy model to generate neural textures 406-408. The UV coordinates represent the axes of the 2D texture. The proxy is used to represent a version of the actual geometry of any or all objects in the class. The architecture 400 can calculate (e.g., generate) a neural texture T for each instance of the object and each represented 3D proxy geometry. i,j =Gen j (w i ), where w i =MLP(z i ) is the latent code z using the MLP NN 404 i Nonlinear reparameterization of .
[0078] Image generators A, B, and C can (e.g., Gen(.)) can represent decoders that receive a latent code (e.g., map (z) 402) as input to generate feature maps, for example, using neural textures 406-410. To render the output view, architecture 400 can rasterize the deferred shadow buffer from each agent, including depth, normal, and UV coordinates. Architecture 400 can then sample the corresponding neural textures 406, 408, and 410, for example, using the shadow buffer UV coordinates of each agent (not shown). The results of the sampling are shown at 414, 416, and 418.
[0079] The architecture 400 can use the contents of the shadow buffer as input to a neural renderer 250 (e.g., a U-Net). The neural renderer 250 can generate four output channels. For example, the neural renderer 250 can generate a color space / color channels 252 representing three output channels (i.e., a red channel, a green channel, and a blue channel). In some embodiments, the color channels 252 can be a color image (e.g., a map) that indicates which colors to render in the image. The fourth output channel can be an alpha channel 254 that represents a mask for a particular object, which specifies how each pixel should be merged with another pixel represented in the object when the two pixels overlap each other. In an example, the alpha channel (e.g., a mask) can represent the opacity of a pair of glasses. That is, the alpha mask can represent the translucency of a particular geometry or surface of an object.
[0080] In some embodiments, multiple neural textures are configured to reconstruct hidden portions of an object captured in the image content. For example, in the view of glasses 406, a portion of the glasses' bow may be hidden because the front view of the glasses hides the bow. The hidden portion (e.g., the bow) can be reconstructed based on a stacked formation (e.g., on top of each other) of neural textures, which can enable a neural renderer to generate (e.g., represent) a transparent layer of the object and a surface behind the transparent layer of the object.
[0081] In some embodiments, the color values can be pre-multiplied by the alpha channel 254 (e.g., mask) because the colors in pixels with low alpha values tend to be particularly noisy in the extracted image mask, which can interfere with the NN 404 (e.g., NN 242). The color channel 252 and the alpha channel 254 can be combined to generate and render the composite image 256.
[0082] In some embodiments, an L1 loss may be calculated by the architecture 400 for both the color channel 252 and the alpha channel 254. In some embodiments, an L1 loss may be calculated by the architecture 400. In some embodiments, a VGG loss may also be calculated for the composite image 256 to account for any perceptual loss in the generated composite image 256.
[0083] In operation, the architecture 400 uses proxy geometry principles to encode geometry using a set of coarse proxy surfaces (e.g., 3D proxy geometry 238) and to encode shape, albedo, and view-dependent effects using a view-dependent neural texture 244. The neural texture 244 is parameterized using a generative model that can generate a variety of shapes and appearances.
[0084] For example, the architecture 400 can generate a neural texture 244 for the 3D proxy geometry 238 generated by the system 200. The 3D proxy geometry 238 typically includes a mesh portion that depicts the geometry and / or texture associated with an object. Using a pose 412 for a particular 3D proxy geometry, the architecture 400 can render a version of the object from a particular viewpoint. For example, normals 426, 428, and 430 are generated as planes representing the object. Depth maps 420, 422, and 424 can also be generated for each pixel of the object. In addition, sampling proxies 414, 416, and 418 can be generated for use as a map in the 3D proxy geometry (e.g., feature map 240) to retrieve specific portions of the geometry for sampling and rendering.
[0085] After generating elements 414-430, architecture 400 can stack the images to generate nine channels, and then can generate multiple views of the object, which can then be concatenated into a deferred shadow buffer. The output of the deferred shadow buffer can be provided to the neural renderer 250, which generates a color space image 252 and an alpha mask.
[0086] In some embodiments, the architecture 400 utilizes a generative latent optimization (GLO) framework to train the NN 404 end-to-end using L1 and VGG perceptual reconstruction losses. In some embodiments, the L1 loss is reconstructed on a composite of pre-multiplied color space channel values, pre-multiplied alpha channel, and a neutral gray background. In some embodiments, the perceptual loss can be applied to the composite image 256, for example, using the second and fifth layers of a VGG pre-trained on a set of images. In some embodiments, the latent code for each class (e.g., map (z) 402) is randomly initialized, and the learning rate of the optimizer is 1e -5 The neural texture 244 (e.g., 406, 408, and 410) can include 9 channels of neural texture. In some embodiments, the map (z) 402 can be represented in 8 dimensions and (w) can be represented in 512 dimensions. For example, the resulting image (e.g., the composite image 256) can be generated for glasses at a resolution of 512x512. Other resolutions can be used for other objects.
[0087] Figures 5A-5CExamples of simulating, capturing, and extracting image content according to embodiments described throughout this disclosure are illustrated. Figure 5A An exemplary apparatus 502 is shown in which an image (e.g., image 504 of a user wearing glasses 506) is captured. While apparatus 502 is depicted as being used to capture glasses objects, other apparatuses can be constructed and used to capture other object categories and use these captures to train neural networks and generate models of object categories. Apparatus 502 depicts a mannequin head simulating a user with a white background and a Calibu calibration configuration to represent a camera and calculate camera geometry and photometric model parameters.
[0088] Figure 5B 5. Image capture using device 502 is shown. Here, four images 508, 510, 512, and 514 are captured to represent multiple poses 412 and an object (e.g., glasses 506). If the object being represented were a car instead of glasses, this step might capture multiple images of the car.
[0089] Figure 5C 4, 516, 518, 520, and 522 are shown, representing possible versions of the glasses. For example, architecture 400 can use images 508-514 to solve for the foreground alpha mask and color. In some embodiments, soft shadows of the glasses (e.g., shadow 524) can be preserved from the matting algorithm. In this example, the latent transform MLP 404 has 4 layers of 256 features, and the rendering U-Net (e.g., neural renderer 250) contains 5 downsampling and upsampling blocks, each with two convolutions (20 convolutions total).
[0090] Figure 6 608 , an image (w) representing a nonlinear latent reparameterization of the latent code (z) 608 , a ground truth image 612 , an exemplary neural texture 614 of the image, and a combined image 616 representing a combined version of the image.
[0091] Figure 6 An example of view interpolation performed by the system described herein compared to ground truth image content according to embodiments described throughout this disclosure is shown. Although the GLO model is generally described above, other view interpolation models can be used, including but not limited to variational autoencoder (VAE) models or game theory (GT) models.
[0092] While providing a specific input angle, other angles of the glasses can be interpolated using a small amount of reconstruction. For example, a left-angle view of the glasses can be provided as input, but the system 200 can reconstruct the view from the right angle by slightly adjusting the input view and reconstructing other viewpoints using neural textures. View-dependent effects captured at the bridge of the glasses can be reconstructed even if they are not captured in the input image.
[0093] System 200 can employ a generative model that allows interpolation in the latent space of an object, effectively constructing a deformable model with a shape and appearance similar to a 3D morphable model. For example, system 200 can generate such an interpolation where the proxy geometry of glasses object 604 remains constant while the latent code (z) 608 is linearly interpolated to generate image (w) 610. The difference can depend on where the model is fitted. The shape of glasses object 604 is realistically shown at image (w) 610 despite the texture mismatch, and an improved overall reconstruction is achieved when all network parameters are fine-tuned.
[0094] Because the system 200 uses a parameterized space of textures, the system can reconstruct a specific instance by finding the correct latent code (z) that reproduces the input view. For example, this can be done by the encoder or by optimizing the reconstruction loss using gradient descent. In some embodiments, the system 200 can alternatively optimize the intermediate parameters of the neural network, including but not limited to optimizing the transformed latent space (w), optimizing the neural texture space, or optimizing all network parameters (i.e., fine-tuning the entire neural network).
[0095] Therefore, given a k} and proxy geometry {P i,1 ,...,P i,K} a set of views {I1,...,I k}, the system 200 can define a new potential code (z) and can set the reconstruction process to optimize as follows:
[0096] z * ,θ * =argmin∑ k 1||I k -Net(z, pk.θ)||1 (Equation 1)
[0097] where Net() is parameterized by the latent code (z), the pose (p), and the intermediate network parameters to be optimized (θ) Figure 4 In some embodiments, the stacked proxy input provides a view of the glasses frame that is occluded by the front proxy, but such a view can be accurately reproduced using the system 200 and architecture 400.
[0098] Figures 7A-7C An exemplary virtual try-on application using the models described herein, according to embodiments described throughout this disclosure, is shown. The generative models utilized by system 200 and architecture 400 can enable a virtual try-on experience for an object. In the depicted example, a user 700 is trying on different glasses 702, 704, and 706, respectively, while being able to move during video / image capture of the user 700 wearing the particular glasses.
[0099] The learned latent space of the glasses (performed by the system 200 and / or architecture 400) can allow the user to modify the appearance and shape of the glasses by modifying the input latent code. Example video image snapshots 708, 710, and 712 illustrate the results of the system 200 processing a video of the user 700 at a close distance without the glasses. For example, the head pose of the user 700 is tracked by the tracking system of the telepresence device 106. The texture proxy can be placed on the head frame of the reference device (e.g., Figure 5A ). The system 200 can then render the neural proxy to generate a color image and alpha mask representing the glasses layers, which can then be composited onto the frame.
[0100] In short, the systems and techniques described herein provide a compact representation for jointly modeling the shape and appearance of objects. The system uses coarse proxy geometry and generative latent textures. The system demonstrates that by jointly modeling a collection of objects, latent interpolation can be performed between visible instances to reconstruct unseen instances with high quality using as few as three input images. The system can assume known 3D proxy geometry and pose.
[0101] Figure 8 is a flow chart illustrating one example of a process 800 for generating a synthetic image based on a 3D proxy geometry model, according to embodiments described throughout this disclosure. Briefly, process 800 can provide an example of using a 3D proxy geometry with a generative model to generate an accurate representation of a 3D object image. Process 800 can utilize at least one processing device and a memory storing instructions that, when executed, cause the processing device to perform the various operations and computer-implemented steps described in the claims. Generally, systems 100, 200, and / or architecture 400 can be used in the description of process 800. In some embodiments, each of systems 100, 200, and architecture 400 can represent a single system.
[0102] At block 802, process 800 includes receiving a gesture associated with an object in image content. In some embodiments, the gesture can be retrieved and / or received based on detecting the object and / or gesture from the image content. For example, process 800 can detect one or more visual cues associated with the object. Visual cues can trigger the detection of a particular object. For example, visual cues can include, but are not limited to, transparent properties, reflective properties, complex geometric shapes, and / or other structural properties captured by the camera that the system 200 determines match a stored category 234 and / or object 236. In some embodiments, the gesture can be evaluated, for example, when glasses are worn on an individual captured by the camera. The gesture can provide knowledge of the location of the user's face, so detection of the glasses can be associated with a location on the face. In some embodiments, process 800 can detect objects at inference time when the task is to replace an object already in the scene with a re-rendered variation of the object.
[0103] For example, the object may be glasses 104" ( Figure 1 ). For example, if user 104 is on a conference call with user 102, glasses 104" may be captured by a camera associated with system 108. Here, the camera may detect glasses 104" and may employ system 200 to generate a realistic view of glasses 104", as conventional capture of glasses 104" may not appear accurately based on reflective and / or transparent surfaces. That is, because an object captured in an image and / or video may include at least a portion of the object's material comprised of a transparent and / or reflective material, process 800 may employ system 200 and / or architecture 400 to correct any representation of the object (glasses 104") to ensure that the object is properly rendered, for example, in 3D, for display to user 102.
[0104] In this example, the image content may include telepresence image data (e.g., as shown in 110) that includes at least a user (e.g., the user in image 104') and the object includes a pair of glasses 104". However, other examples may include image content of other objects having surfaces that are reflective, transparent, and / or otherwise difficult to re-render in a video, for example. In some embodiments, the object includes a portion of a vehicle having reflective properties. For example, when a view of the vehicle portion is re-rendered within a 3D display, the vehicle portion may be reflective and may not appear accurately. In some embodiments, the object includes a portion of any object captured in the image. Thus, process 800 may use generative models, category-level object modeling techniques, and / or other techniques described herein to correct errors and render the content portion.
[0105] At block 804, process 800 includes generating a plurality of three-dimensional (3D) proxy geometries 238 for an object. For example, 3D content modeler 230 can generate 3D proxy geometries 414-430 for glasses 104″, which can represent normal proxy geometries (426, 428, and 430), depth maps (e.g., 420, 422, 424), and sampled versions of the proxies (e.g., 414, 416, and 418). Sample proxies 414, 416, and 418 can represent an atlas (e.g., feature map 240) of geometry and texture samples for particular features of glasses 104″. In some embodiments, each of the plurality of 3D proxy geometries includes a coarse geometric approximation of at least a portion of an object (e.g., glasses 104″) and a potential texture 239 of the object (e.g., glasses 104″) mapped to the coarse geometric approximation (e.g., geometry approximation 246), which can be represented as planes 302, 304, and 306.
[0106] In some embodiments, the plurality of 3D texture proxies encodes a surface light field associated with an object in the image content. For example, the surface light field may include specular reflections associated with the object or other geometric structure reflections away from a particular proxy surface (e.g., lens reflections, refractions, etc.).
[0107] At block 806 , the process 800 includes generating a plurality of neural textures 244 for an object (e.g., glasses 104 ″) based on the plurality of 3D proxy geometries 238 . Here, the neural textures 244 define a plurality of different shapes and appearances representing the object. The neural textures 244 represent at least a portion of the learned feature map 240 that was trained as part of the image capture process. For example, when the glasses object 104 ″ is captured by a camera, the neural texture 244 may be generated using the feature map 240 and the 3D proxy geometries 238 for the object. In operation, the system 200 may generate and store the neural texture 244 for a particular object (or scene) as a mapping onto the 3D proxy geometries 238 for the object.
[0108] At block 808 , the process 800 includes providing the plurality of neural textures 244 to the neural renderer 250 , the plurality of neural textures being provided in a stacked form. For example, the system 200 may use the contents of a shadow buffer (not shown) as input to the neural renderer 250 (eg, a U-Net).
[0109] In operation, the neural renderer 250 can use multiple neural texture inputs to generate an intermediate representation of an object and / or scene to be rendered, for example, using a neural network. The neural texture 244 can be used in conjunction with a 5-layer U-Net, such as the neural network 242 operating with the neural renderer 250, to jointly learn features on the texture map (e.g., feature map 240). The neural renderer 250 can incorporate view-dependent effects using, for example, object-specific convolutional networks to simulate the difference between real appearance (e.g., ground truth) and diffuse reprojection. These effects can be difficult to predict based on scene knowledge, and therefore, a GAN-based loss function can be used to render realistic output.
[0110] In some embodiments, an object (e.g., glasses 104") is associated with a pose (e.g., pose 412). For example, the pose can be a capture angle of an original scene and can be a desired output angle of a composite image that the system 200 and process 800 seek to generate. In such an example, multiple neural textures are based at least in part on the pose. In some embodiments, the neural texture is generated by identifying a category of an object (e.g., glasses) and generating a feature map based on the identified category of the object (e.g., neural texture 244 becomes stacked images 414-430). The feature map can be provided to a neural network 242 (which can be part of a neural renderer / U-net 250). The neural texture 244 can be generated using the feature map 240 based on a view associated with the pose 412. In some embodiments, the neural texture 244 can be generated based on a latent code associated with each instance of the identified category and the view associated with the pose 412.
[0111] In some implementations, the neural renderer uses a generative model to reconstruct unseen object instances within the identified category, and the reconstruction can be based on fewer than four captured views of the object (e.g., glasses 104 ″) (e.g., the three views shown by neural textures 406 , 408 , and 410 ).
[0112] At block 810, the process 800 includes receiving a color image 252 and an alpha mask 254 representing an opacity of at least a portion of an object (glasses 104") from a neural renderer based on a plurality of neural textures. For example, the neural renderer 250 can generate four output channels. That is, the neural renderer 250 can generate a color space color channel 252 representing three output channels (i.e., a red channel, a green channel, and a blue channel). In some implementations, the color image 252 can represent a color space mapping indicating which colors to render for a particular image. The fourth output channel can be An alpha mask 254 representing a channel for a particular object that specifies how each pixel should be merged with another pixel represented in the object when the two pixels overlap. In an example, the alpha mask 254 may represent the opacity of a pair of glasses. In general, the alpha mask 254 may represent the translucency of a particular geometry or surface of an object. In operation, the process 800 may rasterize the neural textures into final image coordinates using, for example, the pose and viewpoint, and may process those textures 252 / 254 into the final image coordinate space of the composite image 256 using a neural renderer.
[0113] At block 812, process 800 includes generating a composite image 256 based on the color image 252 and the alpha mask 256. For example, process 800 may render the latent texture 239 to a target viewpoint (e.g., captured by a camera of system 108). The target viewpoint may be based at least in part on a pose 412 associated with the object (glasses 104"). In some implementations, the 3D texture proxy geometry includes a coarse geometric approximation of at least a portion of the object and a latent texture of the object mapped to the coarse geometric approximation. Although glasses are described in the example of process 800, any number of objects may be substituted and rendered using the techniques of process 800.
[0114] Figure 9An example of a computer device 900 and a mobile computer device 950 that can be used with the techniques described herein is shown. The computing device 900 includes a processor 902, a memory 904, a storage device 906, a high-speed interface 908 connected to the memory 904 and a high-speed expansion port 910, and a low-speed interface 912 connected to a low-speed bus 914 and the storage device 906. Components 902, 904, 906, 908, 910, and 912 are interconnected using various buses and can be mounted on a common motherboard or otherwise installed as appropriate. The processor 902 can process instructions for execution within the computing device 900, including instructions stored in the memory 904 or storage device 906, to display graphical information for a GUI on an external input / output device, such as a display 916 coupled to the high-speed interface 908. In some embodiments, multiple processors and / or multiple buses, as well as multiple memories and memory types, can be used as appropriate. Additionally, multiple computing devices 900 may be connected, with each device providing portions of the necessary operations (eg, as a server bank, a group of blade servers, or a multi-processor system).
[0115] Memory 904 stores information within computing device 900. In one embodiment, memory 904 is one or more volatile memory units. In another embodiment, memory 904 is one or more non-volatile memory units. Memory 904 may also be another form of computer-readable medium, such as a magnetic disk or optical disk.
[0116] Storage device 906 can provide mass storage for computing device 900. In one embodiment, storage device 906 can be or include a computer-readable medium, such as a floppy disk device, a hard disk device, an optical disk device, or a magnetic tape device, a flash memory or other similar solid-state storage device, or a series of devices, including devices in a storage area network or other configuration. A computer program product can be tangibly embodied in an information carrier. A computer program product can also include instructions that, when executed, perform one or more methods, such as those described herein. The information carrier is a computer or machine-readable medium, such as memory 904, storage device 906, or memory on processor 902.
[0117] The high-speed controller 908 manages bandwidth-intensive operations of the computing device 900, while the low-speed controller 912 manages less bandwidth-intensive operations. This allocation of functions is exemplary only. In one embodiment, the high-speed controller 908 is coupled to the memory 904, the display 916 (e.g., via a graphics processor or accelerator), and is coupled to the high-speed expansion port 910 that can accept various expansion cards (not shown). The low-speed controller 912 can be coupled to the storage device 906 and the low-speed expansion port 914. The low-speed expansion port, which can include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), can be coupled to one or more input / output devices such as a keyboard, pointing device, scanner, or network device such as a switch or router through a network adapter.
[0118] As shown, computing device 900 can be implemented in a variety of different forms. For example, it can be implemented as a standard server 920, or multiple times as part of a cluster of such servers. It can also be implemented as part of a rack server system 924. Furthermore, it can also be implemented in a personal computer such as laptop computer 922. Alternatively, components from computing device 900 can be combined with other components in a mobile device (not shown), such as device 950. Each such device can contain one or more computing devices 900, 950, and the entire system can be composed of multiple computing devices 900, 950 communicating with each other.
[0119] Computing device 950 includes, among other components, a processor 952, a memory 964, an input / output device such as a display 954, a communication interface 966, and a transceiver 968. A storage device, such as a microdrive or other device, may also be provided to device 950 to provide additional storage. Each of components 950, 952, 964, 954, 966, and 968 are interconnected using various buses, and multiple components may be mounted on a common motherboard or otherwise as appropriate.
[0120] The processor 952 can execute instructions within the computing device 950, including instructions stored in the memory 964. The processor can be implemented as a chipset including a single or multiple analog and digital processors. For example, the processor can provide coordination of other components of the device 950, such as control of a user interface, applications running on the device 950, and wireless communications of the device 950.
[0121] The processor 952 can communicate with the user through the control interface 958 and the display interface 956 coupled to the display 954. For example, the display 954 can be a TFTLCD (thin film transistor liquid crystal display) or an OLED (organic light emitting diode) display, or other appropriate display technology. The display interface 956 may include appropriate circuits for driving the display 954 to present graphics and other information to the user. The control interface 958 can receive commands from the user and convert them to be submitted to the processor 952. In addition, the external interface 962 can communicate with the processor 952 to enable near-area communication of the device 950 with other devices. For example, in some embodiments using multiple interfaces, the external interface 962 can provide wired or wireless communication.
[0122] Memory 964 stores information within computing device 950. Memory 964 can be implemented as one or more computer-readable media, one or more volatile memory units, or one or more non-volatile memory units. An expansion memory 984 can also be provided and connected to device 950 via an expansion interface 982. For example, expansion interface 982 can include a SIMM (Single In-line Memory Module) card interface. This expansion memory 984 can provide additional storage space for device 950, or it can also store applications or other information for device 950. Specifically, expansion memory 984 can include instructions for executing or supplementing the above-described processes, and can also include security information. Thus, for example, expansion memory 984 can be a security module for device 950 and can be programmed with instructions that allow for secure use of device 950. In addition, secure applications can be provided via a SIMM card and additional information, such as identifying information placed on the SIMM card in an inaccessible manner.
[0123] For example, the memory may include flash memory and / or NVRAM memory, as described below. In one embodiment, a computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more of the methods described above. The information carrier is a computer or machine-readable medium such as memory 964, expansion memory 984, or memory on processor 952, which may be received, for example, via transceiver 968 or external interface 962.
[0124] The device 950 can communicate wirelessly via a communication interface 966, which may include digital signal processing circuitry, as necessary. The communication interface 966 can provide for communication in various modes or protocols, such as GSM voice calls, SMS, EMS or MMS messaging, CDMA, TDMA, PDC, WCDMA, CDMA200, or GPRS, among others. For example, such communication can occur via a radio frequency transceiver 968. In addition, short-range communication can occur, such as using Bluetooth, Wi-Fi, or other such transceivers (not shown). In addition, a GPS (Global Positioning System) receiver module 980 can provide additional navigation and location-related wireless data to the device 950, which can be used appropriately by applications running on the device 950.
[0125] Device 950 may also communicate auditorily using an audio codec 960 that may receive voice information from a user and convert it into usable digital information. Audio codec 960 may also generate audible sounds for the user, such as through a speaker, for example, in an earpiece of device 950. Such sounds may include sounds from voice phone calls, may include recorded sounds (e.g., voice messages, music files, etc.), and may also include sounds generated by applications running on device 950.
[0126] As shown, computing device 950 can be implemented in a variety of different forms. For example, it can be implemented as a cellular phone 980. It can also be implemented as part of a smart phone 982, a personal digital assistant, or other similar mobile device.
[0127] Various implementations of the systems and techniques described herein can be implemented in digital electronic circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system that includes at least one programmable processor, which can be special purpose or general purpose, coupled to a storage system, at least one input device, and at least one output device to receive data and instructions, and coupled to a storage system, at least one input device, and at least one output device to send data and instructions.
[0128] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and may be implemented in high-level procedural and / or object-oriented programming languages, and / or in assembly / machine language. As used herein, the terms "machine-readable medium," "computer-readable medium," and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., a magnetic disk, an optical disk, a memory, a programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0129] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) to display information to the user, and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including sound, voice, or tactile input.
[0130] The systems and techniques described herein can be implemented in a computing system that includes a back-end component (e.g., as a data server), or includes a middleware component (e.g., an application server), or includes a front-end component (e.g., a client computer having a graphical user interface or a web browser through which a user can interact with embodiments of the systems and techniques described herein), or any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), and the Internet.
[0131] A computing system may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship between a client and a server arises through computer programs running on their respective computers and having a client-server relationship to each other.
[0132] In some embodiments, Figure 9 The computing device depicted in FIG may include sensors that interface with a virtual reality headset (VR headset / HMD device 990). For example, Figure 9One or more sensors on the computing device 950 or other computing device depicted in FIG, , can provide input to the VR headset 990, or generally, can provide input to the VR space. Sensors can include, but are not limited to, touch screens, accelerometers, gyroscopes, pressure sensors, biometric sensors, temperature sensors, humidity sensors, and ambient light sensors. The computing device 950 can use the sensors to determine the absolute position and / or detected rotation of the computing device in the VR space, which can then be used as input to the VR space. For example, the computing device 950 can be incorporated into the VR space as a virtual object such as a controller, laser pointer, keyboard, weapon, etc. When a user incorporates the computing device / virtual object into the VR space, its position can allow the user to position the computing device in a certain manner to view the virtual object in the VR space.
[0133] In some embodiments, one or more input devices included in or connected to computing device 950 can be used as input to the VR space. Input devices can include, but are not limited to, a touch screen, a keyboard, one or more buttons, a trackpad, a touchpad, a pointing device, a mouse, a trackball, a joystick, a camera, a microphone, headphones or earbuds with input capabilities, a game controller, or other connectable input devices. When the computing device is incorporated into the VR space, a user interacting with the input devices included on computing device 950 can cause specific actions to occur in the VR space.
[0134] In some embodiments, one or more output devices included on the computing device 950 can provide output and / or feedback to the user of the VR headset 990 in the VR space. The output and feedback can be visual, tactile, or audio. The output and / or feedback can include, but are not limited to, rendering the VR space or virtual environment, vibration, turning on, off, or flashing and / or flashing one or more lights or strobe lights, sounding an alarm, playing a ringtone, playing a song, and playing an audio file. The output device can include, but is not limited to, a vibration motor, a vibration coil, a piezoelectric device, an electrostatic device, a light emitting diode (LED), a strobe light, and a speaker.
[0135] In some embodiments, computing device 950 can be placed within a VR headset 990 to create a VR system. VR headset 990 may include one or more positioning elements that allow computing device 950, such as a smartphone 982, to be placed in an appropriate position within VR headset 990. In such embodiments, the display of smartphone 982 may render a stereoscopic image representing a VR space or virtual environment.
[0136] In some embodiments, computing device 950 can appear as another object in a computer-generated 3D environment. User interactions with computing device 950 (e.g., rotating, shaking, touching a touchscreen, sliding a finger on a touchscreen) can be interpreted as interactions with objects in the VR space. As just one example, computing device 950 can be a laser pointer. In such an example, computing device 950 appears as a virtual laser pointer in the computer-generated 3D environment. As the user manipulates computing device 950, the user in the VR space sees the movement of the laser pointer. The user receives feedback from the interaction with computing device 950 in the VR environment on computing device 950 or VR headset 990.
[0137] In some embodiments, the computing device 950 may include a touch screen. For example, a user can interact with the touch screen in a specific way that can mimic what is happening on the touch screen and what is happening in the VR space. For example, a user can use a pinch-type motion to zoom content displayed on the touch screen. This pinch-type motion on the touch screen can cause the information provided in the VR space to be zoomed. In another example, the computing device can be rendered as a virtual book in a computer-generated 3D environment. In the VR space, pages of the book can be displayed in the VR space, and a swipe of the user's finger across the touch screen can be interpreted as turning / flipping the pages of the virtual book. As each page is turned / flipped, in addition to seeing the page content change, audio feedback can also be provided to the user, such as the sound of turning pages in a book.
[0138] In some embodiments, in addition to the computing device, one or more input devices (e.g., a mouse, a keyboard) may be rendered in the computer-generated 3D environment. The rendered input devices (e.g., a rendered mouse, a rendered keyboard) may be used while rendering in the VR space to control objects in the VR space.
[0139] Computing device 900 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. Computing device 950 is intended to represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are exemplary only and are not meant to limit the disclosed embodiments.
[0140] Furthermore, the logic flows depicted in the accompanying drawings do not require the particular order shown, or sequential order, to achieve the desired results. Furthermore, other steps may be provided or removed from the described flows, and other components may be added to or removed from the described systems. Therefore, other embodiments are within the scope of the appended claims.
Claims
1. A computer-implemented method performed by at least one processing device, the method comprising: receiving a pose associated with an object in the image content; generating a plurality of three-dimensional 3D proxy geometries for the object, the plurality of 3D proxy geometries being based on an appearance of the object; generating a plurality of neural textures for the object based on the plurality of 3D proxy geometries, the plurality of neural textures defining a plurality of different shapes and appearances representing the object; providing the plurality of neural textures to a neural renderer, the plurality of neural textures being provided in a stacked form; receiving, from the neural renderer and based on the plurality of neural textures, a color image and an alpha mask representing opacity of at least a portion of the object; as well as A composite image is generated based on the pose, the color image, and the alpha mask.
2. The method according to claim 1, further comprising: Based at least in part on the pose associated with the object, a latent texture is rendered onto a target viewpoint, wherein each of the plurality of 3D proxy geometries includes a coarse geometric approximation of at least a portion of the object and the latent texture of the object mapped to the coarse geometric approximation.
3. The method according to claim 1, wherein The plurality of neural textures are configured to reconstruct hidden portions of the object captured in the image content, the hidden portions being reconstructed based on the stacked formation of the plurality of neural textures enabling the neural renderer to generate a transparent layer of the object and a surface behind the transparent layer of the object.
4. The method according to claim 1, wherein Each of the plurality of 3D proxy geometries encodes a surface light field associated with the object in the image content, the surface light field including specular reflections associated with the object.
5. The method according to claim 1, wherein The plurality of neural textures are based at least in part on the posture, each neural texture being generated by: identifying the category of the object; generating a feature map based on the identified category of the object; providing the feature map to a neural network; as well as The neural texture is generated based on a latent code associated with each instance of the identified class and a view associated with the gesture.
6. The method according to claim 1, wherein At least a portion of the object is a transparent material.
7. The method according to claim 1, wherein At least a portion of the object is a reflective material.
8. The method according to claim 1, wherein: The image content includes remote presentation image data, the remote presentation image data including at least a user; and The object comprises a pair of glasses.
9. A system for generating content, comprising: at least one processing device; as well as A memory storing instructions that, when executed, cause the system to perform operations including: receiving a pose associated with an object in the image content; generating a plurality of three-dimensional 3D proxy geometries for the object, the plurality of 3D proxy geometries being based on an appearance of the object; generating a plurality of neural textures for the object based on the plurality of 3D proxy geometries, the plurality of neural textures defining a plurality of different shapes and appearances representing the object; providing the plurality of neural textures to a neural renderer, the plurality of neural textures being provided in a stacked form; receiving, from the neural renderer and based on the plurality of neural textures, a color image and an alpha mask representing opacity of at least a portion of the object; as well as A composite image is generated based on the color image and the alpha mask.
10. The system of claim 9, wherein the operations further comprise: Based at least in part on the pose associated with the object, a latent texture is rendered onto a target viewpoint, wherein each of the plurality of 3D proxy geometries includes a coarse geometric approximation of at least a portion of the object and the latent texture of the object mapped to the coarse geometric approximation.
11. The system according to claim 9, wherein: Each of the plurality of 3D proxy geometries encodes a surface light field associated with the object in the image content, the surface light field including specular reflections associated with the object.
12. The system according to claim 9, wherein: The plurality of neural textures are based at least in part on the posture, each neural texture being generated by: identifying the category of the object; generating a feature map based on the identified category of the object; providing the feature map to a neural network; as well as The neural texture is generated based on a latent code associated with each instance of the identified class and a view associated with the gesture.
13. The system according to claim 12, wherein: The neural renderer uses a generative model to reconstruct unseen object instances within the identified category, the reconstruction being based on fewer than four captured views of the object.
14. The system according to claim 9, wherein: The plurality of 3D proxy geometries are based on geometric interpolations that construct shapes of the objects in the image content.
15. A non-transitory machine-readable medium having stored thereon instructions that, when executed by a processor of a computing device, cause the computing device to perform operations comprising: receiving a pose associated with an object in the image content; generating a plurality of three-dimensional 3D proxy geometries for the object, the plurality of 3D proxy geometries being based on an appearance of the object; generating a plurality of neural textures for the object based on the plurality of 3D proxy geometries, the plurality of neural textures defining a plurality of different shapes and appearances representing the object; providing the plurality of neural textures to a neural renderer, the plurality of neural textures being provided in a stacked form; receiving, from the neural renderer and based on the plurality of neural textures, a color image and an alpha mask representing opacity of at least a portion of the object; as well as A composite image is generated based on the color image and the alpha mask.
16. The machine-readable medium of claim 15, wherein the operations further comprise: Based at least in part on the pose associated with the object, a latent texture is rendered onto a target viewpoint, wherein each of the plurality of 3D proxy geometries includes a coarse geometric approximation of at least a portion of the object and the latent texture of the object mapped to the coarse geometric approximation.
17. The machine-readable medium of claim 15, wherein: The plurality of neural textures are configured to reconstruct hidden portions of the object captured in the image content, the hidden portions being reconstructed based on the stacked formation of the plurality of neural textures enabling the neural renderer to generate a transparent layer of the object and a surface behind the transparent layer of the object.
18. The machine-readable medium of claim 15, wherein: The plurality of neural textures are based at least in part on the posture, each neural texture being generated by: identifying the category of the object; generating a feature map based on the identified category of the object; providing the feature map to a neural network; as well as The neural texture is generated based on a latent code associated with each instance of the identified class and a view associated with the gesture.
19. The machine-readable medium of claim 15, wherein: At least a portion of the object is a transparent material.
20. The machine-readable medium of claim 15, wherein: At least a portion of the object is a reflective material.
21. The machine-readable medium of claim 15, wherein: The image content includes remote presentation image data, the remote presentation image data including at least a user; and The object comprises a pair of glasses.
22. The machine-readable medium of claim 15, wherein: The synthetic images are generated using the Generative Latent Optimization (GLO) framework and a perceptual reconstruction loss.
Citation Information
Patent Citations
Systems and methods for estimating pose of textureless objects
CN109074666A
Intelligent naked eye 3D display system based on neural network
CN110324605A