Modeling realistic face wearing glasses

The generative morphable model effectively addresses the challenge of modeling eyewear-facial interactions by jointly capturing geometric and optical interactions, resulting in high-fidelity virtual representations adaptable to different lighting conditions.

CN120322804APending Publication Date: 2025-07-15CTRL-LABS CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380079478.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-29
Filing Date
2023-11-29
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

Existing virtual representation technologies cannot effectively collect geometric and appearance interactions between glasses and faces, resulting in a lack of realism and fidelity in virtual representations, especially when the light changes, the interaction effects between glasses and faces cannot be accurately rendered.

Method used

Generative deformation model is used to model glasses and faces, combined with generative deformation glasses network, and through hybrid grid-volume representation and physical neural lighting, accurate modeling of geometric interactions and photometric interactions between glasses and faces is achieved, supporting the global light transmission effect of different materials, and re-illumination is performed through differentiable neural rendering method.

Benefits of technology

High-fidelity rendering of glasses and faces in a virtual environment can accurately render the interaction between glasses and faces under different lighting conditions, support realism representations of multiple materials, and allow for few samples fitting of new glasses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120322804A_ABST
    Figure CN120322804A_ABST
Patent Text Reader

Abstract

Methods, systems, and storage media for modeling a subject in a virtual environment are disclosed. Exemplary embodiments may: receive image data from a client device, the image data including at least one subject; extracting a face of the at least one subject and an object interacting with the face from the image data, where the object may be glasses worn by the subject; generating a set of facial primitives based on the face, the set of facial primitives including geometry and appearance information; generating a set of object primitives based on the set of potential code of the object; generating an appearance model of luminosity interaction between the face and the object; and rendering the avatar in the virtual environment based on the appearance model, the set of face primitives, and the set of object primitives.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This disclosure relates to U.S. Provisional Application No. 63 / 428,703, filed on Nov. 29, 2022, entitled “VIRTUAL REPRESENTATIONS OF REAL OBJECTS”, and claims the benefit of priority under 35 U.S.C. § 119(e) to this application. For all purposes, the content of this application is hereby incorporated by reference in its entirety. Technical Field

[0003] This disclosure generally relates to improving the virtual representation of real subjects in a virtual environment, and more particularly, to modeling glasses with different materials and the interaction between the face and the glasses to render a more accurate and realistic representation of the face. Background Art

[0004] The virtual representation of users has increasingly focused on social presence and the consequent need to digitize clothing, accessories, and other physical features and characteristics. However, virtual representations cannot effectively capture the geometric and appearance interactions between a person and their accessories. For example, it is a challenge to model the geometric and appearance interactions between glasses and the face to generate a virtual representation (e.g., an avatar) based on that model, and it is often affected by view and time inconsistencies. As a result, they cannot faithfully reconstruct all the geometric and photometric interactions that exist in the real world. Traditional generative models also lack structural priors about the face or glasses, which leads to sub - optimal fidelity. In addition, generative models are not relightable, so we are not allowed to render glasses onto the face with new lighting.

[0005] Therefore, there is a need to provide users with an improved virtual representation that achieves a realistic and lifelike rendering of real objects / accessories that interact with a person in a virtual space. Summary of the Invention

[0006] The present subject matter disclosure provides systems and methods for modeling an avatar of a subject based on a generative morphable model that enables joint modeling of the geometric and photometric interactions of glasses and the face from a dynamic multi - view image set.

[0007] One aspect of the present disclosure relates to a method for modeling a subject in a virtual environment. The method may include: receiving image data from a client device, the image data including at least one subject. The method may include: extracting, from the image data, a face of the at least one subject and an object interacting with the face. The method may include: generating a set of face primitives based on the face, the set of face primitives including geometric structures and appearance information. The method may include: generating a set of object primitives based on a set of latent codes of the object. The method may include: generating an appearance model of photometric interaction between the face and the object. The method may include: rendering an avatar based on the appearance model, the set of face primitives, and the set of object primitives.

[0008] Another aspect of the present disclosure relates to a system configured to model a subject in a virtual environment. The system may include one or more processors configured by machine-readable instructions. The one or more processors may be configured to: receive image data including at least one subject. The one or more processors may be configured to: extract, from the image data, a face of the at least one subject and an object interacting with the face. The one or more processors may be configured to: generate a set of face primitives based on the face, the set of face primitives including geometric structures and appearance information. The one or more processors may be configured to: generate a set of latent codes of the object, the set of latent codes including geometric structure latent codes and appearance latent codes. The one or more processors may be configured to: generate a set of object primitives based on the set of latent codes of the object. The one or more processors may be configured to: generate an appearance model of photometric interaction between the face and the object. The one or more processors may be configured to: render an avatar based on the appearance model, the set of face primitives, and the set of object primitives.

[0009] Yet another aspect of the present disclosure relates to a non-transitory computer-readable storage medium having instructions embodied thereon that are executable by one or more processors to perform a method for modeling a subject in a virtual environment. The method may include: receiving image data from a client device, the image data including at least one subject. The method may include: extracting a face of at least one subject and glasses interacting with the face from the image data. The method may include: generating a set of face primitives based on the face, the set of face primitives including geometric structures and appearance information. The method may include: generating a set of latent codes of the glasses, the set of latent codes including geometric structure latent codes and appearance latent codes. The method may include: generating a set of glasses primitives based on the set of latent codes of the glasses. The method may include: generating an appearance model of photometric interaction between the face and the glasses. The method may include: rendering an avatar based on the appearance model, the set of face primitives, and the set of glasses primitives.

[0010] These and other embodiments will be apparent from the disclosure herein. It should be understood that other configurations of the subject technology will be apparent to those skilled in the art from the following detailed description, wherein various configurations of the subject technology are shown and described by way of example. As will be recognized, the subject technology is capable of other and different configurations, and many details of these configurations can be modified in various other respects, all without departing from the scope of the subject technology. Accordingly, the drawings and detailed description are to be regarded as illustrative in nature and not restrictive. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] For ease of identification of the discussion of any particular element or act, one or more of the most significant digits in the reference numeral refers to the figure number in which the element is first introduced.

[0012] Figure 1 A block diagram showing an overview of a plurality of devices on which some embodiments of the disclosed technology may operate.

[0013] Figure 2 An overview of the overall framework of a generative deformation glasses network according to one or more embodiments.

[0014] Figure 3 According to one or more embodiments Figure 2 The workflow of the deformation geometry module of a generative facial model.

[0015] Figure 4 According to one or more embodiments Figure 2 The workflow of the relighting appearance module of a generative facial model.

[0016] Figures 5A to 5I An illustration of exemplary data acquired for learning according to one or more embodiments.

[0017] Figure 6 Shows glasses replacement on a rendered photorealistic avatar using a generative deformation glasses model according to one or more embodiments.

[0018] Figure 7A And Figure 7B Shows exemplary rendering features according to one or more embodiments.

[0019] Figure 8 Shows the insertion of lenses in a glasses frame included in the rendering of a photorealistic avatar according to one or more embodiments.

[0020] Figure 9 Shows a reconstructed avatar output from a generative deformation glasses network according to one or more embodiments.

[0021] Figure 10 Shows an ablation process regarding geometric guidance according to one or more embodiments.

[0022] Figure 11 Shows an ablation process regarding geometric interaction according to one or more embodiments.

[0023] Figure 12 Shows an ablation process regarding specular reflection features and appearance interaction according to one or more embodiments.

[0024] Figure 13 Shows a comparison of multiple avatars according to one or more embodiments.

[0025] Figure 14 Shows a comparison of multiple avatars according to one or more embodiments.

[0026] Figure 15 Shows a comparison of multiple avatars according to one or more embodiments.

[0027] Figure 16 Shows a block diagram of a system configured to render a model according to certain aspects of the present disclosure.

[0028] Figure 17 Is a flowchart showing multiple steps in a method for rendering a model according to certain aspects of the present disclosure.

[0029] Figure 18 Is a block diagram of an example computer system (e.g., representing both a client and a server) by which aspects of the present subject matter technology can be implemented.

[0030] In one or more embodiments, not all components depicted in each figure are necessary, and one or more embodiments may include additional components not shown in the figures. Variations in the arrangement and type of these components can be made without departing from the scope of the present subject matter disclosure. Additional components, different components, or fewer components can be used within the scope of the present subject matter disclosure. Detailed Description

[0031] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent to one of ordinary skill in the art that some of these specific details may not be required to practice the embodiments of the present disclosure. In other instances, well-known structures and techniques have not been shown in detail so as not to obscure the present disclosure.

[0032] General Overview

[0033] The virtual representation of users has increasingly focused on social presence and the consequent need to digitize clothing, accessories, and other physical features and properties. In particular, glasses play an important role in identity recognition. A realistic virtual representation of the face can greatly benefit from their inclusion. However, virtual representations do not effectively capture the geometric and appearance interactions between a person and accessories such as glasses. Glasses and the face deform each other's geometric structures at their points of contact. Similarly, appearance is coupled through global light transport, and shadows and mutual reflections can occur and affect radiance, and appearance changes can also be caused by light transport. Therefore, traditional models for modeling the geometric and appearance interactions between glasses and the face of a person's virtual representation cannot capture these physical interactions because they model glasses and the face independently.

[0034] Traditional real-time graphics engines support the synthesis of individual components (e.g., hair, clothing), but the interactions between the face and other objects must necessarily be approximated with overly simplified physics-inspired constraints or heuristics (e.g., "no interpenetration"). Therefore, they cannot faithfully reconstruct all the geometric and photometric interactions that exist in the real world and are generally not performed in real time with physically based rendering. Others have attempted to solve the interaction as a 2D image synthesis problem but have been plagued by view and temporal inconsistencies. The lack of 3D information including contact and occlusion leads to limited fidelity and incoherent results in moving and varying views. Generative modeling for the face and glasses may be able to cover the shape and appearance variations of each object class, but the interactions between the objects are not considered in these methods, resulting in unreasonable object synthesis.

[0035] In addition, traditional generative models also lack structural priors about the face or glasses, which leads to suboptimal fidelity and the inability to model realistic interactions. Moreover, generative models cannot re-light, thus preventing rendering glasses on the face with new lighting. Standard representations cannot handle such diverse materials that exhibit significant transmission effects, and inferring the parameters of these materials for realistic re-lighting remains challenging.

[0036] The embodiments disclosed herein describe methods and systems for providing solutions rooted in computer technology, namely, 3D deformation and relightable glasses models that represent the shape and appearance of a glasses frame and the interaction of the glasses frame with the face to generate a photorealistic virtual representation (e.g., an avatar). Instead of modeling the glasses in isolation, the model takes into account the glasses and their interaction with the face to achieve photorealism. Aspects of the embodiments model the geometric and photometric interactions between the glasses frame and the face, accurately combining high-fidelity geometric and photometric interaction effects to better capture the representation of the glasses interacting with the face.

[0037] According to embodiments, the model can include explicit volumetric primitives that move and deform to effectively achieve expressive animations with semantic correspondences across the frame. Different from traditional volumetric methods, the model naturally preserves the correspondences across the glasses, thus greatly simplifying explicit modifications to the geometry, such as lens insertion and frame deformation.

[0038] In some embodiments, to effectively support large changes in the glasses topology, a hybrid representation that combines surface geometry and volumetric representation is adopted. Adopting the hybrid representation provides explicit correspondences across the glasses, and thus, the model can slightly deform the structure of the glasses based on the head shape.

[0039] In some embodiments, the model can be conditioned by a high-fidelity generative human head model, allowing the model to specifically handle the deformations and appearance changes caused by wearing glasses. In some embodiments, the model can include a deformed facial model based on glasses-conditioned deformation and an appearance network that combines the interaction effects caused by wearing glasses. In some implementations, the glasses frame is modeled without lenses to avoid modeling lens refraction.

[0040] According to embodiments, the methods incorporate physics-based neural lighting into generative modeling. These methods infer the output radiance given a view, point light source position, visibility, and specular reflections with multiple lobe sizes. The generative model is relightable under point light sources and natural lighting. This significantly improves the generalization ability and supports subsurface scattering, reflection, and high-fidelity rendering of various materials (including translucent plastics and metals) in a single model. Embodiments model global light transport effects (e.g., cast shadows between the face and the glasses) and can adapt to new glasses through inverse rendering.

[0041] In some embodiments, as a pre - processing step, a differentiable neural signal distance function (SDF) can be used, for example, to separately reconstruct the geometry of the glasses from multi - view images. The regularization term based on these pre - computed geometries of the glasses significantly improves the fidelity of the model.

[0042] Example architecture

[0043] Figure 1 is a schematic diagram of an environment 100 in which the methods, apparatuses, and systems described herein according to various embodiments can be implemented. As Figure 1 shown, the environment 100 may include a user device 110, a platform 120, and a network 130. The devices in the environment 100 may be interconnected by a wired connection, a wireless connection, or a combination of a wired connection and a wireless connection.

[0044] The user device 110 includes one or more devices capable of receiving, generating, storing, processing, and / or providing information associated with the platform 120. For example, the user device 110 may include a computing device (e.g., a desktop computer, a laptop computer, a tablet computer, a handheld computer, a smart speaker, a server, etc.), a mobile phone (e.g., a smartphone, a wireless phone, etc.), a headset, or other wearable devices (e.g., a virtual reality or augmented reality headset, smart glasses, a smart watch), or similar devices. In some embodiments, the user device 110 may receive information from the platform 120 and / or send information to the platform 120 via the network 130.

[0045] The platform 120 includes one or more devices as described elsewhere herein. In some embodiments, the platform 120 may include a cloud server or a group of cloud servers. In some embodiments, the platform 120 may be designed to be modular such that software components can be swapped in or out. Thus, the platform 120 can be easily and / or quickly reconfigured for different uses.

[0046] In some embodiments, as shown, the platform 120 may be hosted in a cloud computing environment 122. It is noted that although the various embodiments described herein describe the platform 120 as being hosted in the cloud computing environment 122, in some embodiments, the platform 120 may not be cloud - based (i.e., may be implemented outside of a cloud computing environment) or may be partially cloud - based.

[0047] The cloud computing environment 122 includes the environment hosting the platform 120. The cloud computing environment 122 can provide computing services, software services, data access services, data storage (e.g., database) services, etc., which do not require the end user (e.g., the user device 110) to know the physical location and configuration of one or more systems and / or one or more devices of the hosting platform 120. As shown, the cloud computing environment 122 can include a set of computing resources 124 (collectively referred to as "computing resources 124" and individually referred to as "computing resource 124").

[0048] The computing resources 124 include one or more personal computers, one or more workstation computers, one or more server devices or other types of computing devices and / or communication devices. In some embodiments, the computing resources 124 can host the platform 120. The computing resources 124 can include an application programming interface (API) layer that controls the various applications in the user device 110. The API layer can also provide a tutorial to the user of the user device 110 about the new features in the application. The cloud resources can include computing instances executed in the computing resources 124, storage devices provided in the computing resources 124, data transmission devices provided by the computing resources 124, etc. In some embodiments, the computing resources 124 can communicate with other computing resources 124 through a wired connection, a wireless connection, or a combination of a wired connection and a wireless connection.

[0049] As Figure 1 Further shown, the computing resources 124 include a set of cloud resources, such as one or more applications ("APP") 124-1, one or more virtual machines ("VM") 124-2, one or more virtualized storage devices ("VS") 124-3, or one or more hypervisors ("HYP") 124-4, etc.

[0050] Application 124-1 includes one or more software applications that can be provided to and / or accessed by user device 110 and / or platform 120. Application 124-1 can eliminate the need to install and execute software applications on user device 110. For example, Application 124-1 can include software associated with platform 120 and / or any other software that can be provided via cloud computing environment 122. In some embodiments, one Application 124-1 can send information to or receive information from one or more other Applications 124-1 via virtual machine 124-2. Application 124-1 can include one or more modules configured to perform operations in accordance with aspects of the embodiments. These modules will be described in detail later.

[0051] Virtual machine 124-2 includes a software implementation of a machine (e.g., a computer) that executes programs like a physical machine. Virtual machine 124-2 can be a system virtual machine or a process virtual machine, depending on how much and to what extent virtual machine 124-2 uses and corresponds to any real machine. A system virtual machine can provide a complete system platform that supports the execution of a complete operating system ("OS"). A process virtual machine can execute a single program and can support a single process. In some embodiments, virtual machine 124-2 can execute on behalf of a user (e.g., user device 110) and can manage the infrastructure of cloud computing environment 122, such as data management, synchronization, or long-duration data transfer.

[0052] Virtualized storage device 124-3 includes one or more of the following storage systems and / or one or more devices: the one or more storage systems and / or one or more devices that use virtualization techniques in the storage system or storage device of computing resources 124. Virtualized storage device 124-3 can include storage instructions that, when executed by a processor, cause computing resources 124 to perform at least some of the operations in one or more methods consistent with the present disclosure. In some embodiments, in the context of a storage system, the types of virtualization can include block virtualization and file virtualization. Block virtualization can refer to the extraction (or separation) of logical storage from physical storage, such that the storage system can be accessed without regard to the physical storage or heterogeneous structure. This separation can allow the administrator of the storage system to have flexibility in how the administrator manages the storage of end users. File virtualization can eliminate the dependence between the data accessed at the file level and the location of the physical storage of the file. This can enable optimized storage usage, server consolidation, and / or performance of non-disruptive file migration.

[0053] The hypervisor 124-4 can provide hardware virtualization technologies that allow multiple operating systems (e.g., "guest operating systems") to execute simultaneously on a host computer (e.g., computing resource 124). The hypervisor 124-4 can present a virtual operating platform to the guest operating systems and can manage the execution of the guest operating systems. Multiple instances of multiple operating systems can share the virtualized hardware resources.

[0054] The network 130 can include, for example, any one or more of a local area network (LAN), a wide area network (WAN), and the Internet, etc. In addition, the network 130 can include, but is not limited to, any one or more of the following network topologies: these network topologies include bus networks, star networks, ring networks, mesh networks, star-bus networks, and tree or hierarchical networks, etc.

[0055] Figure 1 The number and arrangement of the devices and networks shown are provided as an example. In fact, there can be more devices and / or networks, fewer devices and / or networks, devices and / or networks different from those shown Figure 1 in the figure, or devices and / or networks with different arrangements than those shown in the figure. In addition, Figure 1 two or more of the devices shown in the figure can be implemented in a single device; or Figure 1 a single device shown in the figure can be implemented as multiple distributed devices. Additionally or alternatively, a set of devices (e.g., one or more devices) in the environment 100 can perform one or more functions described as being performed by another set of devices in the environment 100. Figure 1 The present disclosure provides systems and methods for modeling faces and glasses using a generative deformation glasses network. The network includes a deformation geometry component and a relightable appearance component for providing a generative model of glasses and faces and their interactions, which can be converted for use in a virtual environment, for example, through an avatar / virtual representation. Embodiments are capable of jointly modeling the geometric and photometric interactions of glasses and faces from a collection of dynamic multi-view images.

[0056] The various embodiments described herein address the above and other drawbacks by providing a method for correctly modeling glasses with different materials and the interactions between faces and glasses using a generative model. The generative model of glasses uses a hybrid mesh-volume representation to represent the topological changing shapes and complex appearances of glasses. A physics-inspired neural relighting method can also be implemented to support the global light transport effects of different materials in a single model.

[0057]

[0058] Figure 2 This is an overview of the overall framework of the Generative Deformable Glasses Network 200 for rendering models according to one or more embodiments. As Figure 2 shown, the Generative Deformable Glasses Model 210 of the network 200 can consider two main components (i.e., the Deformable Geometry Module 202 and the Re-lightable Appearance Module 204) to support the replacement of glasses on the face.

[0059] By training the Generative Deformable Glasses Model 210 based on one or more datasets 220 (including various faces under different illuminations and glasses on faces with and without glasses), the Generative Deformable Glasses Model 210 enables smooth glasses replacement, directional light rendering, environment map rendering, and lens insertion for few-shot reconstruction of images. The Generative Deformable Glasses Network 200 can generate new (generative) glasses through latent code modification and support the replacement of re-lightable materials while maintaining the shape.

[0060] The Deformable Geometry Module 202 separately learns the generative face model and the generative glasses model to model the changes of the face and glasses and the geometric interaction between the face and glasses, so that these models can be combined together. The Re-lightable Appearance Module 204 accurately renders the re-lightable appearance by calculating the features representing the light interaction with the re-lightable face model to allow the joint re-lighting of the face and glasses. The Re-lightable Appearance Module 204 correctly models glasses with different materials and the interaction between the face and glasses.

[0061] The Lens Insertion Module 206 can be configured to implement lens insertion with attractive lens reflection and refraction effects. Once the Generative Deformable Glasses Model 210 is trained, it can reconstruct and re-light the invisible glasses with only a small amount of input. According to the embodiments, the Generative Deformable Glasses Network 200 preserves the correspondence between the primitives. Therefore, it is easy to solve the problem of inserting lenses in the generated glasses by selecting the control points of the lens contour on a single template. In some embodiments, the Lens Insertion Module 206 is also configured to combine physically accurate refraction and reflection based on the glasses prescription.

[0062] The Rendering Module 208 can be configured to render a realistic synthesis of the reconstructed glasses on the head of the volumetric avatar from any viewpoint under the new illumination generated based on the Generative Deformable Glasses Model 210 and output the reconstructed image 230. The Generative Deformable Glasses Network 200 supports a differentiable rendering model (or avatar), enabling few-shot reconstruction from a small number of perspective images through inverse rendering. The non-re-lightable appearance model and the re-lightable appearance model share the same latent code. Therefore, the Rendering Module 208 only uses the full illumination from the new illumination to render the few-shot reconstruction.

[0063] Figure 3 For a generative face model according to one or more embodiments Figure 2 detailed workflow of the deformation geometry module 202 of the generative face model.

[0064] According to various embodiments, the generative deformation glasses model is based on a mixture of volumetric primitives (MVP), which is a unique volumetric neural rendering method for achieving real-time high-fidelity rendering. The generative deformation glasses model includes explicit volumetric primitives that move and deform to effectively achieve expressive animations with cross-frame semantic correspondences. Different from grid-based methods, the generative deformation glasses model supports topological changes in geometric structures.

[0065] To model a face without glasses, the generative face model employs a pre-trained face encoder 302 and a face decoder 304. Given the encoding of a facial expression facial identity encoding of geometric structure and facial identity encoding of texture then the facial primitive geometry and appearance 306 are decoded into:

[0066]

[0067] wherein, is the face decoder 304, G f ={t, R, s} is the position of the facial primitive rotation and scale tuple; is the opacity of the facial primitive; is the RGB color of the facial primitive in the fully illuminated image. N fprim represents the number of facial primitives, and M represents the resolution of each primitive among multiple primitives. In some embodiments, N fprim is set to be equal to 128×128, and M is set to be equal to 8.

[0068] To model glasses, the generative glasses model includes a variational autoencoder (hereinafter referred to as "glasses encoder 308") and a glasses decoder 310. The glasses encoder 308 can be encoded according to the following equation:

[0069]

[0070] wherein, ε g is the glasses encoder 308, which takes the one-hot vector of the glasses as input and generates the geometric structure latent code of the glasses and appearance latent codes as output. These latent codes are decoded at the glasses decoder 310 to generate the glasses primitives 312 as follows:

[0071]

[0072] where is the glasses decoder 310, G g ={t g , R g , s g} is a tuple of the position, rotation, and scale of the glasses primitive, where the position rotation scale is the opacity of the glasses primitive; is the RGB color of the glasses primitive in the fully lit image. N gprim represents the number of glasses primitives. In some embodiments, N fprim is set to be equal to 32×32. Thus, the face and glasses are modeled in a separate latent space.

[0073] The deformations caused by the interaction between the glasses and the face are modeled as residual deformations of the primitives. The interaction 314 calculates the residual deformations G δf , G δg as follows:

[0074]

[0075] where G δf ={δ t , δ R , δ s} and are the residuals of the position, rotation, and scale from their values in the canonical space. The interaction can affect the glasses in two different ways: non-rigid deformations caused by fitting to the head; rigid deformations caused by facial expressions. According to the embodiments, these two effects are modeled separately to better generalize to new combinations of glasses and identities. The deformation residuals can be modeled as:

[0076]

[0077] where uses the facial identity information to deform the glasses to fit the target head, and Take the facial expression encoding as input to model the relative rigid motion of glasses on the face caused by different facial expressions (e.g., upward sliding when wrinkling the nose). The synthesis 316 synthesizes the generative facial model, the generative glasses model, and the interaction between the face and the face together (i.e., G f +G δf 、O f 、C f model the face; and G g +G δg 、O g 、C g model the glasses) to generate a synthetic face 318 wearing glasses.

[0078] Figure 4 For the detailed workflow of the relightable appearance module 204 of the generative facial model according to one or more embodiments. Figure 2

[0079] According to the embodiments, the appearance values of the primitives under the uniform tracking illumination of the deformation geometry module 202 (i.e., C f and C g ) are used to learn the geometric structure and the deformation caused by the interaction. According to the embodiments, the relightable appearance module 204 models the relightable face in the relightable appearance model by introducing a relightable appearance decoder 402, so as to achieve relighting of the generative facial model. Train the relightable appearance decoder 402 with the encoding of facial expressions and facial textures , and additionally, the relightable appearance decoder 402 decodes and generates the appearance A f conditioned on the line-of-sight direction v and the light direction l as:

[0080]

[0081] where, is the relightable appearance decoder 402, is the appearance board composed of RGB colors under a single point light source.

[0082] To model the photometric interaction of glasses on the face, the photometric interaction is regarded as a residual regulated by the glasses latent code, similar to the deformation residual. Calculate the light features representing the light interaction, such as shadow features and specular reflection features. Casting shadows is the most obvious appearance (light) interaction of glasses on the face. Therefore, the generative facial model includes a shadow encoder 404, which takes the shadow features as input to facilitate shadow modeling according to the following formula:

[0083]

[0084] Among them, is the shadow encoder 404, is the appearance residual of the face; is the shadow feature calculated by accumulating the opacity when the light steps to the primitive from each of the multiple light sources, representing light visibility. Therefore, the shadow feature represents the first reflection of light transmitted to the face and the glasses.

[0085] The appearance of the relightable glasses is modeled similarly to the relightable face. According to various embodiments, for the purpose of modeling the glasses on the face, the appearance of the relightable glasses can be defined as a conditional model with facial elements, such that the occlusion and multiple reflections of light by the avatar's head have been incorporated into the appearance. With facial expressions Encoding of facial texture The color C of the glass primitive g The relightable glasses decoder 406 is trained, and additionally, the relightable glasses decoder 406 decodes conditional on the line-of-sight direction v, the light direction l, the shadow feature, and the specular reflection feature, and generates the glasses appearance A g as:

[0086]

[0087] Among them, is the glasses appearance board, and is the specular reflection feature; A shadow is the shadow feature calculated according to Equation (8), and Equation (8) encodes the facial information.

[0088] The compositor 408 generates the photorealistic avatar 410 by compositing the appearance and residual of the face (i.e., A f +A δf ) with the glasses appearance A g , which models the geometric and photometric interactions between the glasses and the face under fully relit conditions.

[0089] In some embodiments, using a bidirectional reflectance distribution function (BRDF) parameterized as a spherical Gaussian function with three different lobes, the specular reflection feature A spec at each point on the primitive is calculated based on the normal, the light direction, and the line-of-sight direction. Explicitly adjusting the specular reflection significantly improves the fidelity of relighting and enhances the generalization ability to various frame materials.

[0090] According to various embodiments, differentiable volume rendering can be used to render the predicted volume primitives. The positions of all the primitives in space can be represented as G. Thus, when rendering only the face without any glasses, the positions of the face primitives will be represented as G = G f , and when rendering the face with glasses, the positions of the face primitives will be represented as G = {G f + G δf , G g + G δg}. Similarly, the opacity of all the primitives in space can be represented as O, which will be in the form of O = O f or O = {O f , O g} for the non - glasses - wearing and glasses - wearing renderings respectively. The color of all the primitives in space can be represented as C, where, in the fully - lit image, C = C f , and C = {C f , C g}, while in the relit frame, C = A f and C = {A f + A δf , A g}. Then, differentiable volume aggregation is used to render the image based on the face G, opacity O, and color primitives C respectively.

[0091] Figures 5A to 5I FIG. is an illustration of exemplary data acquired for learning according to one or more embodiments. The generative deformable glasses model is a generative model designed to learn about glasses, faces, and the interaction between glasses and faces. Thus, the generative deformable glasses model collects three types of data in the dataset: glasses, faces, and faces with glasses.

[0092] In some embodiments, the lenses can be removed from the glasses for all the datasets to decouple the learning framework style from the lens effects (which vary according to prescriptions).

[0093] In some embodiments, a set of glasses can be selected from the dataset to cover various sizes, styles, and materials, including metals and translucent plastics in various colors. Multiple multi - perspective images (e.g., 70 multi - perspective images) can be acquired for each glasses instance using an acquisition system (e.g., user device 110). In some implementations, a multi - perspective light - table acquisition system (e.g., a camera) is used to acquire the dataset. Figure 5A Shows an image of glasses 502 and multi - perspective images 504 and 506 of glasses 502. Key points are identified for each of the multiple glasses (e.g., glasses 502 and multi - perspective images 504, 506). Figure 5B Shows the key - point detection 508 of glasses 502.

[0094] Extract a three-dimensional (3D) mesh for each of the plurality of glasses using, for example, a surface reconstruction method. Figure 5C An exemplary glasses mesh 510 extracted based on glasses 502 is shown. The 3D mesh will later provide supervision for the glasses MVP geometry. However, geometric changes occur once the glasses are worn. Accordingly, embodiments implement Bounded Biharmonic Weight (BBW) to define a rough deformation model that is used to adapt the mesh to the glasses-wearing face dataset using keypoint detection 508.

[0095] According to embodiments, a dataset of faces without glasses and the same set of faces with glasses is collected. The dataset can include a set of subjects (e.g., 25 subjects). Figure 5D A set of images 514 collected using a multi-view light stage acquisition system with a camera is shown, the set of images 514 including a subject without glasses. For example, the subject can be instructed to perform various facial expressions, resulting in recordings with varying expressions and head poses. Each subject can be collected a specified number of times. For example, each subject is collected three times: a first frame without glasses and second and third frames with glasses randomly selected from a set of glasses. These frames can be collected at different camera viewpoints. Figure 5E A set of images 516 of a subject wearing randomly selected glasses is shown.

[0096] To allow relighting, data is collected under different lighting conditions. Accordingly, the acquisition system uses time-multiplexed lighting. A fully illuminated frame (i.e., a frame with all light sources on the light stage turned on) can be inserted every three frames to allow tracking, and the remaining two-thirds of the frames are used to observe the subject under varying lighting conditions in which only a subset of the plurality of light sources (“group” light sources) are turned on, as Figure 5F shown, Figure 5F A set of images 518 in which only a subset of the plurality of light sources are turned on is shown.

[0097] According to embodiments, the data is preprocessed using a multi-view face tracker to generate a rough but topologically consistent face mesh for each frame. For example, a first face mesh, a second face mesh, and a third face mesh can be generated corresponding to the first frame, the second frame, and the third frame, respectively. Figure 5G A face mesh 524 generated based on frames of a set of collected images 514 is shown. Tracking and detection are performed on the fully illuminated frames and inserted into the partially illuminated frames as necessary. For the set of glasses-wearing face portions of the dataset, a set of keypoints (e.g., 20 keypoints) are detected on the glasses. Figure 5HShows a set of faces 520 wearing glasses (e.g., a set of faces 516 wearing glasses), where the glasses include a set of key points. As Figure 5I shown, a face and glasses segmentation mask 522 is also generated. The set of key points on the glasses of the set of faces wearing glasses (i.e., 520) and the face and glasses segmentation mask 522 are used to fit a glasses BBW mesh deformation model to match the observed glasses.

[0098] Figure 6 Shows the glasses replacement on a rendered photorealistic avatar using a generative deformable glasses model according to one or more embodiments. The generative deformable glasses model generates new glasses through latent code modification, and thus, supports changing the material 610 and shape 620 of the glasses while adjusting to fit different faces / heads. According to various embodiments, the glasses can be seamlessly replaced or substituted with other re-lightable materials while maintaining the shape.

[0099] Figure 7A and Figure 7B Shows exemplary rendering features according to one or more embodiments. As Figure 7A shown, various embodiments may include directional light rendering 702 of the avatar (e.g., group lighting). As Figure 7A shown, various embodiments may include environment map rendering 704.

[0100] Figure 8 Shows the insertion of the lens 820 of the glasses frame included in the rendering of a photorealistic avatar 804 according to one or more embodiments. In some embodiments, the lens 820 may include prescription-based refractive and reflective features.

[0101] Figure 9 Shows a reconstructed avatar output from a generative deformable glasses network based on an input image 900 according to one or more embodiments. The output includes a first view 902 of the reconstructed avatar from different camera viewpoints, a second view 904 of the reconstructed avatar, and a third view 906 of the reconstructed avatar.

[0102] According to various embodiments, the generative deformable glasses network 200 can be trained in the following two stages: deformation geometry training for training the deformation geometry module 202; and re-lightable appearance training for training the re-lightable appearance module 204. In the first stage, fully illuminated images are used to train the geometry of the face and glasses. In the second stage, images under group light sources are used to train the re-lightable appearance model.

[0103] According to various embodiments, the deformation geometry training uses the following equation to optimize the expression encoder ε f , the glasses encoder ε g and the decoder G δf, G δf The parameter Φ g :

[0104]

[0105] wherein, the parameter Φ g is optimized on N I different subjects; are different fully illuminated frames including with and without glasses; N C are different camera viewpoints; I i represents all the ground truth camera images and associated processed assets of a frame, including facial geometry, glasses geometry, face segmentation, and glasses segmentation; similarly, I r represents the reconstructed images and corresponding assets from volume rendering. The full illumination loss function consists of three main parts:

[0106]

[0107] wherein, is the photometric reconstruction loss, which is defined as:

[0108]

[0109] wherein, is the l1 loss between the observed image and the reconstructed image; are the VGG loss and the GAN loss. is the geometry-guided loss, which is calculated based on the separately reconstructed glasses (e.g., glasses 502 and the multi-view images 504, 506 of glasses 502) to improve the geometric accuracy of the glasses, so as to better separate the face and the glasses in the joint training. The joint training can be defined as:

[0110]

[0111] This equation includes the chamfer distance loss the glasses mask loss and the glasses segmentation loss These losses prompt the network 200 to separate the identity-dependent deformations from the glasses-intrinsic deformations, thus helping the network to generalize across different identities.

[0112] In some embodiments, the following regularization loss is introduced in the first training stage

[0113]

[0114] wherein, is the KL divergence loss between the prior Gaussian distribution and the distribution of the glasses latent space; The l2 norm is used to suppress the incremental (delta) deformation of the face to reduce the large displacement of face primitives.

[0115] During the training process, the weight of each loss term can be set to λ L1 = 1, λ vgg = 1, λ gan = 1, λ c = 0.01, λ m = 10, λ s = 10, λ KL = 10 -4 , λ L2 = 10 -3 .

[0116] According to various embodiments, once the deformation geometry module 202 is trained, the parameter Φ g is frozen, and the relit appearance training phase begins to train the appearance A f , A δf , A g . The appearance A f , A δf , A g 's parameters are denoted as Φ a . The relit appearance training uses the following equation to optimize the parameter Φ a :

[0117]

[0118] where the parameter Φ a is optimized on N I different subjects; is a different set of lighting frames including wearing glasses and not wearing glasses on the face; N C is different camera viewpoints.

[0119] In some embodiments, for frames illuminated by a group of light sources, the two closest fully illuminated frames can be used to use G f , G g , G δf , G δg and linearly interpolate to generate the face and glasses geometry to obtain the face and glasses geometry for the group light source image. The objective function of the second stage is the mean squared error photometric loss The VGG loss and the GAN loss are not used for relit appearance training because doing so will introduce blocky artifacts in the reconstruction.

[0120] Figure 10Illustrates ablation process 1000 for geometric guidance according to some embodiments. Process 1000 shows losses for geometric guidance including surface normals and segmentation, which are essential for achieving clear and sharp glasses reconstruction.

[0121] Images 1001a, 1001b, 1002a, and 1002b are created using input image 1003a, and image 1003b is a magnified image of glasses input image 1003a. A model without geometric guidance is trained only with image-based reconstruction and regularization losses. As Figure 10 shown, without geometric guidance, the reconstructed avatar in image 1001a fails to reconstruct the detailed geometry of the glasses. The magnified image of the glasses in image 1001b shows a lack of detail and clarity in the nose pads. Without geometric guidance, the reconstruction results in a blurred outcome. In contrast, the generative deformable glasses model with geometric guidance (hereinafter referred to as the "full method") achieves higher geometric fidelity, thus generating an avatar with clear and accurate glasses in image 1002a. The magnified image of the glasses in image 1002b shows details in the nose pads.

[0122] Table 1 summarizes the quantitative ablation results of process 1000 for each part of the generative deformable glasses model. The table shows that the full method produces the best reconstruction / output.

[0123] Table 1: Quantitative Ablation Results

[0124]

[0125] Figure 11 Illustrates ablation process 1100 for geometric interaction according to some embodiments. When the glasses and the face interact, the glasses and the face deform each other at the contact points. Process 1100 shows that without modeling this deformation, aspects of the rendering are inaccurate. When geometric interaction is modeled, the generative deformable glasses model learns and faithfully represents the deformation of the head and the nose.

[0126] Figure 11As shown, the reconstructed avatar 1101a without deformation causes the glasses to be blocked by the nose. The magnified image 1101b of the glasses shows that a part of the glasses seems to be blocked by the nose and almost seems to be inside the nose. The reconstructed avatar 1102a without deformation shows that the temple of the glasses is rendered incorrectly and penetrates into the head. The magnified image 1102b of the glasses shows that a part of the glasses seems to be blocked by the subject's hair. In contrast, the generative deformation glasses model with deformation generates avatars 1103a and 1104a with accurate interactive rendering, effectively modeling facial deformation. The magnified images 1103b and 1104b of the glasses show proper and accurate rendering at the interaction points between the face / head and the glasses.

[0127] In some embodiments, the generative deformation glasses network includes generating deformation heatmaps 1105 and 1106 to illustrate the deformation caused by the interaction between the glasses and the face.

[0128] Figure 12 An ablation process 1200 regarding specular reflection features and appearance interaction is shown. Process 1200 shows the physics-inspired features of neural relighting, where the proposed specular reflection (A spec ) and shadow features (A shadow ) are effective on neural relighting. Images 1201, 1202, 1203, 1204, 1205, and 1206 are created using input images 1207 and 1209, and image 1208 is a magnified image of the glasses from input image 1207. As Figure 12 shown, the rendering without using the specular reflection feature 1201 cannot reconstruct the specular highlights on the frame highlighted in the magnified image 1202. Additionally, as shown in image 1203, the model without appearance interaction cannot reconstruct the correct shadows on the face. In contrast, as shown in images 1204 and 1205, the model with specular reflection features shows more details in the glasses, and the model with appearance interaction 1206 shows the correct lighting and shadows on the face.

[0129] Table 1 includes a summary of the quantitative ablation results of process 1200 for the test and evaluation regarding the protruding frame. The table shows that the full method produces the best reconstruction / output. Adding each part effectively improves the performance of all metrics.

[0130] Figure 13Shows a comparison of avatars 1301, 1302, 1303, and 1304 generated based on input images 1305 and 1306. Avatars 1301 and 1302 are generated using the Generative Latent Textured Objects (GeLaTO) model, and avatars 1303 and 1304 are generated using the generative deformable glasses model of the various embodiments. GeLaTO is capable of generative modeling of glasses. However, it assumes that everything in the scene except the glasses is static. GeLaTO was re-implemented and trained using the dataset described in Figures 5A to 5I with ground truth masks. Since GeLaTO does not support relighting, only fully illuminated frames are compared in Figure 13 .

[0131] As Figure 13 shown, due to the billboard-based geometry of GeLaTO, the rendering of avatars 1301 and 1302 lacks geometric details and generates incorrect occlusion boundaries. In contrast, the generative deformable glasses model achieves high-fidelity results and generates correct occlusions in 3D. Table 2 summarizes the quantitative results of the generative deformable glasses model and shows that the model significantly outperforms other models in all metrics.

[0132] Table 2: Quantitative Comparison Using GeLaTO

[0133]

[0134] Figure 14 Shows a comparison of avatars 1402-1, 1402-2, 1402-3, and 1402-4 (collectively referred to hereinafter as "avatars 1402") created using GIRAFFE with avatars 1404-1, 1404-2, 1404-3, and 1404-4 (collectively referred to hereinafter as "avatars 1404") created using a model based on generative adversarial network for video editing (VideoEdit-GAN-based model). Compared with avatars 1406-1, 1406-2, 1406-3, and 1406-4 (collectively referred to hereinafter as "avatars 1406") created using the generative deformable glasses model, the results of GIRAFFE and the VideoEdit-GAN-based model cannot render Figure One consistent results.

[0135] GIRAFFE includes a compositional neural radiance field that supports adding and changing objects in the scene. However, the official implementation only supports objects within the same category. Figure 14It shows that unsupervised synthetic generative modeling still results in sub-optimal fidelity with limited resolution. VideoEdit-GAN is a state-of-the-art (SOTA) image-based editing method that allows users to insert glasses onto a facial image. As Figure 14 shown, image-based methods cannot maintain color and view consistency. In addition, this method cannot select specific types of glasses. In contrast, the models according to various embodiments can accurately reproduce glasses and faces with a rendering that is consistent in view and time.

[0136] Figure 15 Shows a comparison of relighting in an avatar 1501 created using the Lumos model based on input image 1503 and an avatar 1502 created using the generative deformable glasses model based on input image 1503. Referring above Figure 13 and Figure 14 All the methods mentioned do not support relighting of faces and glasses. Therefore, the relighting results of Lumos are compared with the models according to various embodiments. Lumos is a SOTA method for portrait relighting in image space. As Figure 15 shown, due to the lack of 3D information, Lumos has difficulty rendering non-local light transport effects such as shadows cast by glasses. In contrast, the models according to various embodiments generate reasonable soft shadows and more accurately model the photometric interaction between the face and the glasses.

[0137] Therefore, the generative deformable glasses model is a 3D deformation and relightable model that is used to create realistic glasses synthesis on a volumetric avatar head from any viewpoint under new lighting. Experiments show (e.g., as shown in Figures 7 to Figure 15 shown) that it is possible to reproduce geometric and photometric interactions in the real world by leveraging neural rendering with a hybrid mesh-volumetric generative model. By explicitly controlling the motion of primitives, various embodiments achieve a learning-based modeling of the geometric interaction between glasses and faces. Various embodiments (e.g., as Figure 12 shown) also illustrate the effectiveness of using physics-inspired lighting features as inputs for neural relighting and demonstrate that the generative deformable glasses model can relight multiple materials that are both transmissive and reflective using a single generative model. Finally, the generative deformable glasses model allows for few-shot fitting of new glasses, thus allowing relighting without additional online learning and training data. In some embodiments, the network can perform few-shot fitting on real-scene images by adopting a method similar to fine-tuning at test time or through physically accurate fitting of the lens achieved by inverse rendering.

[0138] Figure 16FIG. 1600 is a block diagram of a system 1600 configured to render a model in accordance with certain aspects of the present disclosure. In some embodiments, system 1600 may include one or more computing platforms 1602 (e.g., may correspond to platform 120). The one or more computing platforms 1602 may be configured to communicate with one or more remote platforms (e.g., computing resources 124) according to a client / server (e.g., user device 110 / cloud computing environment 122) architecture, a peer-to-peer architecture, and / or other architectures. The one or more remote platforms 1604 may be configured to communicate with other remote platforms through the one or more computing platforms 1602 and / or according to a client / server architecture, a peer-to-peer architecture, and / or other architectures. A user may access system 1600 through the one or more remote platforms 1604 or a client.

[0139] The one or more computing platforms 1602 may be configured by machine-readable instructions 1606. The machine-readable instructions 1606 may include one or more instruction modules. The instruction modules may include computer program modules. The instruction modules may include one or more of the following modules: a receiving module 1608; an extraction module 1610; a facial primitive generation module 1612; an object primitive generation module 1614; a deformation recognition module 1616; a specular reflection feature calculation module 1618; a shadow feature calculation module 1620; an appearance generation module 1622; a rendering module 1624; and / or other instruction modules.

[0140] The receiving module 1608 may be configured to receive image data from a client device. The first device may include a VR, MR, AR device, a camera, or a mobile phone, etc. The first device may run a VR / AR application. The image data may include an image of a subject or a scene including at least one subject.

[0141] The extraction module 1610 may be configured to identify and extract at least one subject's face and an object interacting with the face from the image data. The object may be glasses worn by the subject. The extraction module 1610 may also be configured to identify and extract the subject's head from the image data.

[0142] The facial primitive generation module 1612 may be configured to generate a set of facial primitives based on the subject's face. The set of facial primitives may include the geometric structure and appearance information of the face. In some embodiments, the set of facial primitives may further include the geometric structure and appearance information of the subject's head. To generate the set of facial primitives, the facial primitive generation module 1612 may also be configured to decode the encoding of facial expressions, facial geometry, and facial texture, where the set of facial primitives includes: a tuple of the position, rotation, and scale of the set of facial primitives; the opacity of the facial primitives; and the color of the facial primitives.

[0143] The object primitive generation module 1614 can be configured to generate a set of potential codes for an object. This set of potential codes can include the geometric structure potential code and the appearance potential code of the object. A set of object primitives is generated based on this set of potential codes of the object. To generate this set of object primitives, the object primitive generation module 1614 can also be configured to decode this set of potential codes of the object. This set of object primitives can include: a tuple of the position, rotation, and scale of this set of object primitives; the opacity of the object primitive; and the color of the object primitive.

[0144] The deformation recognition module 1616 can be configured to recognize the deformations caused by the interaction between the object and the face. These deformations include two different deformation types: non-rigid deformations caused by the object adapting to the subject and rigid deformations caused by the facial expressions of the subject.

[0145] The deformation recognition module 1616 can also be configured to model the deformation as the residual deformation of this set of facial primitives and this set of object primitives. The residual deformation for handling these two deformation types can be generated by deforming the geometric structure and appearance of the object to adapt to the face (or head) of the subject and inferring the movement of the object on the face. The deformation can be based on the geometric structure potential code of the object and the facial identity information. The movement of the object on the face can be the result of the movement caused by the facial expressions of the subject when wearing the object.

[0146] The specular reflection feature calculation module 1618 can be configured to calculate the specular reflection features at each point on this set of object primitives.

[0147] The shadow feature calculation module 1620 can be configured to calculate the shadow features representing the first reflection of light transmitted to the face and the object.

[0148] The appearance generation module 1622 can be configured to generate an appearance model of the photometric interaction between the face and the object. This appearance model includes a relightable appearance face model. This relightable appearance face model can be generated based on this set of facial primitives, more specifically, based on the facial expression information, facial texture information, and the color of the facial primitives. This appearance model includes a relightable appearance object model. This relightable appearance object model can be generated based on this set of object primitives, more specifically, based on the object texture information, object identity information, specular reflection features, and shadow features.

[0149] The appearance generation module 1622 can also be configured to model the photometric interaction of the object on the face based on the appearance residual. The appearance generation module 1622 can also be configured to determine the appearance residual of the face based on the shadow features, light direction, facial texture information in this set of facial primitives, and object texture information in this set of object primitives. The line-of-sight direction and light direction can be identified based on the image data.

[0150] The rendering module 1624 can be configured to render an avatar of the subject based on the appearance model, the set of facial primitives, and the set of object primitives, while adjusting the avatar based on the deformation and photometric interactions between the face and the object. The avatar can be a virtual 3D, photorealistic representation of the subject. The avatar can be rendered in a virtual environment at a second device. The second device can be a VR / AR headset or the like. In some embodiments, the first device and the second device can be the same or different client devices.

[0151] In some embodiments, one or more computing platforms 1602, one or more remote platforms 1604, and / or external resources 1626 can be operably linked via one or more electronic communication links. For example, such electronic communication links can be established at least in part via a network such as the Internet and / or other networks. It will be appreciated that this is not intended to be limiting, and the scope of the present disclosure includes embodiments in which one or more computing platforms 1602, one or more remote platforms 1604, and / or external resources 1626 can be operably linked via some other communication medium.

[0152] A given remote platform 1604 can include one or more processors configured to execute computer program modules. These computer program modules can be configured to enable an expert or user associated with the given remote platform 1604 to interact with the system 1600 and / or external resources 1626, and / or to provide other functions attributed herein to one or more remote platforms 1604. By way of non-limiting example, a given remote platform 1604 and / or a given computing platform 1602 can include one or more of the following: a server; a desktop computer; a laptop computer; a handheld computer; a tablet computing platform; a netbook; a smartphone; a game console; and / or other computing platforms.

[0153] External resources 1626 can include information sources external to the system 1600, external entities participating in the system 1600, and / or other resources. In some embodiments, some or all of the functions attributed herein to external resources 1626 can be provided by resources included in the system 1600.

[0154] One or more computing platforms 1602 can include an electronic storage device 1628, one or more processors 1630, and / or other components. One or more computing platforms 1602 can include communication lines or ports to enable information exchange with a network and / or other computing platforms. In Figure 16The illustration of one or more computing platforms 1602 is not intended to be limiting. One or more computing platforms 1602 can include multiple hardware components, software components, and / or firmware components that operate together to provide the functionality attributed herein to one or more computing platforms 1602. For example, one or more computing platforms 1602 can be implemented by a cloud of computing platforms that operate together as one or more computing platforms 1602.

[0155] The electronic storage device 1628 can include a non-transitory storage medium that stores information electronically. The electronic storage medium of the electronic storage device 1628 can include one or both of a system storage device and a removable storage device, where the system storage device is configured to be integrated with one or more computing platforms 1602 (i.e., substantially non-removable), and the removable storage device can be removably connected to one or more computing platforms 1602 via, for example, a port (e.g., a USB port, a FireWire port, etc.) or a drive (e.g., a disk drive, etc.). The electronic storage device 1628 can include one or more of the following: an optically readable storage medium (e.g., an optical disc, etc.); a magnetically readable storage medium (e.g., a magnetic tape, a magnetic hard disk drive, a floppy disk drive, etc.); a charge-based storage medium (e.g., an electrically erasable programmable read-only memory (EEPROM), a random access memory (RAM), etc.); a solid-state storage medium (e.g., a flash drive, etc.); and / or other electronically readable storage media. The electronic storage device 1628 can include one or more virtual storage resources (e.g., a cloud storage device, a virtual private network, and / or other virtual storage resources). The electronic storage device 1628 can store software algorithms, information determined by one or more processors 1630, information received from one or more computing platforms 1602, information received from one or more remote platforms 1604, and / or other information that enables one or more computing platforms 1602 to operate as described herein.

[0156] One or more processors 1630 can be configured to provide information processing capabilities in one or more computing platforms 1602. Accordingly, one or more processors 1630 can include one or more of the following: a digital processor; an analog processor; a digital circuit designed to process information; an analog circuit designed to process information; a state machine; and / or other mechanisms for electronically processing information. Although one or more processors 1630 are Figure 16is shown as a single entity in the figures, but this is for illustrative purposes only. In some embodiments, one or more processors 1630 may include multiple processing units. These processing units may be physically located within the same device, or one or more processors 1630 may represent the processing functionality of multiple devices operating in concert. One or more processors 1630 may be configured to execute modules 1608, 1610, 1612, 1614, 1616, 1618, 1620, 1622, and / or 1624, and / or other modules. One or more processors 1630 may be configured to execute modules 1608, 1610, 1612, 1614, 1616, 1618, 1620, 1622, and / or 1624, and / or other modules by: software; hardware; firmware; some combination of software, hardware, and / or firmware; and / or other mechanisms for configuring processing capabilities on one or more processors 1630. As used herein, the term "module" may refer to any component or set of components that performs the functions attributed to the module. This may include one or more physical processors, processor-readable instructions, circuits, hardware, storage media, or any other component during execution of processor-readable instructions.

[0157] It should be appreciated that although modules 1608, 1610, 1612, 1614, 1616, 1618, 1620, 1622, and / or 1624 are shown in Figure 16is shown as being implemented within a single processing unit, but in embodiments in which one or more processors 1630 include multiple processing units, one or more of the modules 1608, 1610, 1612, 1614, 1616, 1618, 1620, 1622, and / or 1624 may be implemented remotely from other modules. The descriptions of the functions provided for the different modules 1608, 1610, 1612, 1614, 1616, 1618, 1620, 1622, and / or 1624 below are for illustrative purposes and are not intended to be limiting, since any of the modules 1608, 1610, 1612, 1614, 1616, 1618, 1620, 1622, and / or 1624 may provide more or fewer functions than those described. For example, one or more of the modules 1608, 1610, 1612, 1614, 1616, 1618, 1620, 1622, and / or 1624 may be excluded, and some or all of its functions may be provided by other modules among the modules 1608, 1610, 1612, 1614, 1616, 1618, 1620, 1622, and / or 1624. As another example, one or more processors 1630 may be configured to execute one or more additional modules, which may perform some or all of the functions attributed to one of the modules 1608, 1610, 1612, 1614, 1616, 1618, 1620, 1622, and / or 1624 below.

[0158] The techniques described herein: may be implemented as one or more methods executed by one or more physical computing devices; may be implemented as one or more non-transitory computer-readable storage media storing multiple instructions that, when executed by one or more computing devices, cause the execution of the one or more methods; or may be implemented as one or more physical computing devices specially configured with a combination of hardware and software to execute the one or more methods.

[0159] Figure 17 is a flowchart showing multiple steps in a method 1700 for rendering a model according to certain aspects of the present disclosure. For purposes of explanation, reference is made herein to Figures 1 to 16 describe an example method 1700. Further for purposes of explanation, the steps of the example method 1700 are described herein as being performed in sequence or linearly. However, multiple instances of the example method 1700 may be performed in parallel.

[0160] In step 1702, method 1700 includes: receiving image data from a client device, the image data including at least one subject. In step 1704, method 1700 includes: extracting a face of the at least one subject and an object interacting with the face from the image data. In step 1706, method 1700 includes: generating a set of face primitives based on the face, the set of face primitives including geometric structures and appearance information. In step 1708, method 1700 includes: generating a set of latent codes for the object, the set of latent codes including a geometric latent code and an appearance latent code. In step 1710, method 1700 includes: generating a set of object primitives based on the set of latent codes for the object. In step 1712, method 1700 includes: identifying deformations caused by the interaction of the object with the face. In step 1714, method 1700: generates an appearance model of the photometric interaction between the face and the object. In step 1716, method 1700: renders an avatar of the subject based on the appearance model, the set of face primitives, the set of object primitives, and the deformation.

[0161] According to one aspect, the object is glasses worn by the subject.

[0162] According to one aspect, generating a set of face primitives includes: decoding encodings of facial expressions, facial geometry, and facial textures, wherein the set of face primitives includes: a tuple of the position, rotation, and scale of the set of face primitives; the opacity of the face primitives; and the color of the face primitives.

[0163] According to one aspect, generating a set of object primitives includes: decoding the set of latent codes for the object, wherein the set of object primitives includes: a tuple of the position, rotation, and scale of the set of object primitives; the opacity of the object primitives; and the color of the object primitives.

[0164] According to one aspect, method 1700 may include: modeling the deformation as a residual deformation of the set of face primitives and the set of object primitives. These deformations may include non-rigid deformations caused by fitting the object to the subject and rigid deformations caused by the facial expressions of the subject.

[0165] According to one aspect, method 1700 may include: deforming the geometry and appearance of the object to fit the face or head of the subject based on the geometric latent code for the object and facial identity information; inferring the movement of the object on the face based on the facial expressions of the subject; and generating a deformation residual based on the deformation and the movement.

[0166] According to one aspect, the appearance model includes: a relightable appearance facial modeling based on facial expression information, facial texture information in the set of facial primitives, and colors of the facial primitives in the set of facial primitives; and a relightable appearance object modeling based on object texture information in the set of object primitives, object identity information in the set of object primitives, specular reflection features, and shadow features.

[0167] According to one aspect, method 1700 may include: calculating specular reflection features at each point on the set of object primitives.

[0168] According to one aspect, method 1700 may include: calculating shadow features that represent the first reflection of light transmitted onto the face and the object.

[0169] According to one aspect, method 1700 may include: identifying a line-of-sight direction and a light direction based on image data.

[0170] According to one aspect, method 1700 may include: determining an appearance residual of the face based on shadow features, light direction, facial texture information in the set of facial primitives, and object texture information in the set of object primitives.

[0171] Figure 18 FIG. is a block diagram showing an exemplary computer system 1800 by which aspects of the present subject matter can be implemented. In some aspects, computer system 1800 can be implemented using hardware, or a combination of software and hardware, which can be either in a dedicated server, integrated into another entity, or distributed across multiple entities.

[0172] Computer system 1800 (e.g., a server and / or a client) includes a bus 1808 or other communication mechanism for conveying information, and a processor 1802 coupled to bus 1808 for processing information. As an example, computer system 1800 can be implemented using one or more processors 1802. Processor 1802 can be a general-purpose microprocessor, a microcontroller, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a programmable logic device (PLD), a controller, a state machine, gate logic, discrete hardware components, or any other suitable entity that can perform computations or other information operations.

[0173] In addition to hardware, computer system 1800 may also include code that creates an execution environment for the computer programs being discussed. For example, code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. This code is stored in the included memory 1804, which may be, for example, Random Access Memory (RAM), flash memory, Read-Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), registers, a hard disk, a removable disk, a Compact Disc Read-Only Memory (CD-ROM), a Digital Versatile Disc (DVD), or any other suitable storage device. The memory 1802 is coupled to the bus 1808 for storing information and instructions to be executed by the processor 1802. The processor 1802 and the memory 1804 may be supplemented by, or incorporated into, dedicated logic circuitry.

[0174] Instructions may be stored in the memory 1804 and may be implemented in one or more computer program products, i.e., in one or more modules of computer program instructions encoded on a computer-readable medium for execution by, or to control the operation of, the computer system 1800, and in accordance with any method known to those skilled in the art, these computer program instructions include, but are not limited to, the following computer languages: data-oriented languages (e.g., SQL, dBase), system languages (e.g., C, Objective-C, C++, Assembly), structural languages (e.g., Java,.NET), and application languages (e.g., PHP, Ruby, Perl, Python). The instructions may also be implemented in the following computer languages: such as array languages, aspect-oriented languages, assembly languages, authoring languages, command-line interface languages, compiled languages, concurrent languages, curly-bracket languages, dataflow languages, data-structured languages, declarative languages, esoteric languages, extension languages, fourth-generation languages, functional languages, interactive-mode languages, interpreted languages, iterative languages, list-based languages, little languages, logic-based languages, machine languages, macro languages, meta-programming languages, multi-paradigm languages, numerical analysis, non-English-based languages, class-based object-oriented languages, prototype-based object-oriented languages, off-side rule languages, procedural languages, reflective languages, rule-based languages, scripting languages, stack-based languages, synchronous languages, syntax handling languages, visual languages, wirth languages, and xml-based languages. The memory 1804 may also be used to store temporary variables or other intermediate information during the execution of instructions to be executed by the processor 1802.

[0175] A computer program as discussed herein does not necessarily correspond to a file in a file system. The program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple cooperating files (e.g., files that store one or more modules, subroutines, or portions of code). The computer program can be deployed to execute on one computer or multiple computers, which are located at one site or distributed across multiple sites and interconnected by a communication network. The processes and logical flows described in this specification can be performed by one or more programmable processors that execute one or more computer programs to perform functions by operating on input data and generating output.

[0176] The computer system 1800 also includes a data storage device 1806, such as a magnetic disk or optical disk, coupled to the bus 1808 for storing information and instructions. The computer system 1800 can be coupled to various devices through an input / output module 1810. The input / output module 1810 can be any input / output module. Exemplary input / output module 1810 includes a data port such as a USB port. The input / output module 1810 is configured to connect to a communication module 1812. Exemplary communication module 1812 includes a network interface card, such as an Ethernet card and a modem. In some aspects, the input / output module 1810 is configured to connect to multiple devices, such as an input device 1814 and / or an output device 1816. Exemplary input device 1814 includes a keyboard and a pointing device, such as a mouse or a trackball, through which a user can provide input to the computer system 1800. Other types of input devices 1814 can also be used to provide interaction with the user, such as a haptic input device, a visual input device, an audio input device, or a brain-computer interface device. For example, the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or haptic feedback, and the input from the user can be received in any form, including acoustic input, voice input, haptic input, or brainwave input. Exemplary output device 1816 includes a display device for displaying information to the user, such as a liquid crystal display (LCD) monitor.

[0177] In accordance with one aspect of the present disclosure, the above-described game system can be implemented using computer system 1800 in response to one or more sequences of instructions included in memory 1804 being executed by processor 1802. Such instructions can be read into memory 1804 from another machine-readable medium (e.g., data storage device 1806). Execution of the sequence of instructions contained in main memory 1804 causes processor 1802 to perform the process steps described herein. One or more processors in a multiprocessing arrangement can also be employed to execute the sequence of instructions contained in memory 1804. In an alternative aspect, hardwired circuitry can be used in place of or in combination with software instructions to implement various aspects of the present disclosure. Accordingly, aspects of the present disclosure are not limited to any particular combination of hardware circuitry and software.

[0178] Aspects of the subject matter described in this specification can be implemented in a computing system that includes a backend component (e.g., a data server), or that includes a middleware component (e.g., an application server), or that includes a frontend component (e.g., a client computer having a graphical user interface or a web browser through which a user can interact with an implementation of the subject matter described in this specification); or aspects of the subject matter described in this specification can be implemented in any combination of one or more such backend, middleware, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). The communication network can include, for example, any one or more of the following: a LAN; a WAN; and the Internet, among others. Additionally, the communication network can include, but is not limited to, for example, any one or more of the following network topologies, including a bus network, a star network, a ring network, a mesh network, a star-bus network, or a tree or hierarchical network, among others. The communication module can be, for example, a modem or an Ethernet card.

[0179] Computer system 1800 includes a client and a server. The client and the server are typically remote from each other and typically interact through a communication network. The relationship of the client and the server is generated by computer programs running on respective computers and having a client-server relationship with each other. Computer system 1800 can be, for example, but is not limited to: a desktop computer, a laptop computer, or a tablet computer. Computer system 1800 can also be embedded in another device, which can be, for example, but is not limited to: a mobile phone, a personal digital assistant (PDA), a mobile audio player, a global positioning system (GPS) receiver, a video game console, and / or a television set-top box.

[0180] As used herein, the term "machine-readable storage medium" or "computer-readable medium" refers to any one or more media that participate in providing instructions to processor 1802 for execution. Such media may take many forms, including but not limited to non-volatile media, volatile media, and transmission media. For example, non-volatile media includes optical or magnetic disks, such as data storage device 1806. Volatile media includes dynamic memory, such as memory 1804. Transmission media includes coaxial cables, copper wire, and fiber optics, including the wires that make up bus 1808. Common forms of machine-readable media include, for example, a floppy disk, flexible disk, hard disk, magnetic tape, any other magnetic medium, CD-ROM, DVD, any other optical medium, punch cards, paper tape, any other physical medium with patterns of holes, RAM, PROM, EPROM, FLASH EPROM, any other memory chip or cartridge, or any other medium readable by a computer. A machine-readable storage medium may be a machine-readable storage device, a machine-readable storage substrate, a storage device, a combination of substances that affect a machine-readable propagated signal, or a combination of one or more of them.

[0181] When a user's computing system 1800 reads game data and provides a game, information can be read from the game data and stored in a storage device (e.g., memory 1804). Additionally, data from a memory 1804 server accessed via a network, bus 1808, or data storage device 1806 can be read and loaded into memory 1804. Although the data is described as being found in memory 1804, it will be understood that the data need not be stored in memory 1804 and can be stored in other memories (e.g., data storage device 1806) accessible to processor 1802 or distributed across several media.

[0182] As used herein, the phrase "at least one of" after a series of items, together with the terms "and" or "or" used to separate any of the items, modifies the list as a whole, rather than modifying each member of the list (i.e., each item). The phrase "at least one of" does not require the selection of at least one item; rather, the phrase allows the meaning that it includes at least one of any one of the items, and / or at least one of any combination of the items, and / or at least one of each of the items. As an example, the phrase "at least one of A, B, and C" or "at least one of A, B, or C" both refer to: only A, only B, or only C; any combination of A, B, and C; and / or at least one of each of A, B, and C.

[0183] As used in this specification or in the claims, terms such as "comprising" or "having" are intended to be open-ended in a manner similar to the way the term "including" is construed as "including" when used as a transitional word in a claim. The word "exemplary" as used herein means "serving as an example, instance, or illustration". Any embodiment described herein as "exemplary" should not necessarily be construed as more preferred or advantageous than other embodiments.

[0184] Unless otherwise specified, a reference to an element in the singular is not intended to mean "one and only one" but rather "one or more". All structural and functional equivalents of the elements of the various configurations described throughout this disclosure that are known or later become known to those of ordinary skill in the art are expressly incorporated herein by reference and are intended to be encompassed by the claimed subject matter. In addition, nothing disclosed herein is intended to be dedicated to the public, whether or not such disclosure is expressly recited in the foregoing description.

[0185] Although this specification contains many details, these details should not be construed as limitations on the scope that may be claimed, but rather as descriptions of particular embodiments of the subject matter. Certain features that are described in the context of different embodiments in the present invention may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. In addition, although the features may have been described above as acting in certain combinations and even initially claimed as such, in some cases, one or more features from the claimed combination may be removed from the combination, and the claimed combination may be directed to a sub-combination or a variation of a sub-combination.

[0186] The subject matter of the present invention has been described in terms of specific aspects, but other aspects may be implemented and are within the scope of the appended claims. For example, although the operations are depicted in the figures in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in a sequential order, nor should it be construed as requiring that all of the shown operations be performed to achieve the desired result. The acts recited in the claims may be performed in a different order and still achieve the desired result. As an example, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of the various system components in the aspects described above should not be understood to require such separation in all aspects, but rather it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products. Other variations are within the scope of the appended claims.

Claims

1. A computer-implemented method for modeling a subject in a virtual environment, the computer-implemented method being executed by at least one processor, the method comprising: Receiving image data from a client device, the image data including at least one subject; Extracting, from the image data, a face of the at least one subject and an object interacting with the face; Generating a set of face primitives based on the face, the set of face primitives including geometric structures and appearance information; Generating a set of object primitives based on a set of latent codes of the object; Generating an appearance model of the photometric interaction between the face and the object; And Rendering an avatar based on the appearance model, the set of face primitives, and the set of object primitives.

2. The computer-implemented method according to claim 1, wherein, The object includes glasses worn by the at least one subject.

3. The computer-implemented method according to claim 1, further comprising: Generating the set of latent codes of the object, the set of latent codes including geometric structure latent codes and appearance latent codes.

4. The computer-implemented method according to claim 1, further comprising: Decoding encodings of facial expressions, facial geometry, and facial textures, wherein the set of face primitives includes: a tuple of positions, rotations, and scales of the set of face primitives; the opacity of the face primitives; and the colors of the face primitives; and Generating the set of face primitives based on the decoding.

5. The computer-implemented method according to claim 1, further comprising: Decoding the set of latent codes of the object to generate the set of object primitives, wherein the set of object primitives includes: a tuple of positions, rotations, and scales of the set of object primitives; the opacity of the object primitives; and the colors of the object primitives.

6. The computer-implemented method according to claim 1, further comprising: Identifying deformations caused by the interaction of the object with the face; And Modeling the deformations as residual deformations of the set of face primitives and the set of object primitives.

7. The computer-implemented method according to claim 6, wherein, The deformations include non-rigid deformations caused by the object adapting to the at least one subject and rigid deformations caused by facial expressions of the at least one subject.

8. The computer-implemented method according to claim 1, further comprising: Deforming the geometry and appearance of the object based on the geometric structure latent code of the object and facial identity information to adapt to the face or head of the at least one subject; Inferring the movement of the object on the face based on the facial expressions of the at least one subject; And Generating a deformation residual based on the deformations and the movement.

9. The computer-implemented method according to claim 1, wherein, The appearance model includes: re-lightable appearance face modeling, the re-lightable appearance face modeling being based on facial expression information, facial texture information, and the colors of the face primitives in the set of face primitives; and re-lightable appearance object modeling, the re-lightable appearance object modeling being based on object texture information in the set of object primitives, object identity information in the set of object primitives, specular reflection features, and shadow features.

10. The computer-implemented method according to claim 1, further comprising: Calculating specular reflection features at each point on the set of object primitives.

11. The computer-implemented method according to claim 1, further comprising: Calculating shadow features representing the first reflection of light transmitted to the face and the object; Identifying a line-of-sight direction and a light direction based on the image data; And Determining an appearance residual of the face based on the shadow features, the light direction, the facial texture information in the set of facial primitives, and the object texture information in the set of object primitives.

12. A system, comprising: One or more processors; And A memory storing instructions that, when executed by the one or more processors, cause the system to: Receive image data including at least one subject; Extract a face of the at least one subject and an object interacting with the face from the image data; Generate a set of facial primitives based on the face, the set of facial primitives including geometric structure and appearance information; Generate a set of latent codes for the object, the set of latent codes including a geometric structure latent code and an appearance latent code; Generate a set of object primitives based on the set of latent codes of the object; Generate an appearance model of the photometric interaction between the face and the object; And Render an avatar based on the appearance model, the set of facial primitives, and the set of object primitives.

13. The system according to claim 12, wherein, The one or more processors are further configured to: Decode encodings of facial expressions, facial geometry, and facial textures, wherein the set of facial primitives includes: a tuple of position, rotation, and scale of the set of facial primitives; opacity of the facial primitives; and color of the facial primitives; and Generate the set of facial primitives based on the decoding result.

14. The system according to claim 12, wherein The one or more processors are further configured to: Decode the set of latent codes of the object to generate the set of object primitives, wherein the set of object primitives includes: a tuple of position, rotation, and scale of the set of object primitives; opacity of the object primitives; and color of the object primitives; and Generate the set of object primitives based on the decoding result.

15. The system according to claim 12, wherein, The one or more processors are further configured to: Identify deformations caused by the interaction of the object with the face; and Model the deformations as residual deformations of the set of facial primitives and the set of object primitives.

16. The system according to claim 15, wherein, The deformations include non-rigid deformations caused by the object adapting to the at least one subject and rigid deformations caused by facial expressions of the at least one subject.

17. The system according to claim 12, wherein The one or more processors are further configured to: Deform the geometry and appearance of the object to adapt to the face or head of the at least one subject based on the geometric structure latent code of the object and facial identity information; Infer the movement of the object on the face based on the facial expressions of the at least one subject; And Generate a deformation residual based on the deformed geometry and appearance of the object and the movement.

18. The system according to claim 12, wherein The appearance model includes: relightable appearance facial modeling based on facial expression information, facial texture information in the set of facial primitives, and the color of the facial primitives in the set of facial primitives; and relightable appearance object modeling based on object texture information in the set of object primitives, object identity information in the set of object primitives, specular reflection features, and shadow features.

19. The system according to claim 12, wherein, The one or more processors are further configured to: calculate specular reflection features at each point on the set of object primitives; calculate shadow features representing the first reflection of light transmitted onto the face and the object; identify the line-of-sight direction and the light direction based on the image data; and determine the appearance residual of the face based on the shadow features, the light direction, facial texture information in the set of facial primitives, and object texture information in the set of object primitives.

20. A non-transitory computer-readable storage medium having instructions embodied thereon that are executable by one or more processors to perform a method for modeling a subject in a virtual environment, the method comprising: receiving image data from a client device, the image data including at least one subject; extracting a face of the at least one subject and glasses interacting with the face from the image data; generating a set of facial primitives based on the face, the set of facial primitives including geometric structures and appearance information; generating a set of latent codes for the glasses, the set of latent codes including geometric structure latent codes and appearance latent codes; generating a set of glasses primitives based on the set of latent codes for the glasses; generating an appearance model for photometric interaction between the face and the glasses; and rendering an avatar based on the appearance model, the set of facial primitives, and the set of glasses primitives.