Avatar animation with generic pre-trained facial motion coding
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2026-04-10
Smart Images

Figure CN121844357A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates in general to systems and techniques for generating three-dimensional (3D) models. For example, aspects of this disclosure relate to avatar animation using general pre-trained facial motion codes for faces. Background Technology
[0002] Many devices and systems allow a scene to be captured by generating frames (also called images) and / or video data (including multiple images or frames). For example, a camera or a computing device that includes a camera (e.g., a mobile device (such as a mobile phone or smartphone) that includes one or more cameras) can capture a sequence of frames of a scene. Frame and / or video data can be captured and processed by such devices and systems (e.g., mobile devices, IP cameras, etc.) and can be output for consumption (e.g., displayed on that device and / or other devices). In some cases, frame and / or video data can be captured by such devices and systems and output for processing and / or consumption by other devices.
[0003] Frames can be processed (e.g., using object detection, recognition, segmentation, etc.) to determine the objects present in the frame, which is useful for many applications. For example, a model can be determined to represent the objects in the frame, and this model can be used to facilitate the efficient operation of various systems. Examples of such applications and systems include augmented reality (AR), robotics, automotive and aerospace, 3D scene understanding, object grasping, object tracking, and many other applications and systems. Summary of the Invention
[0004] This paper describes systems and techniques for generating textured three-dimensional (3D) facial models. In an exemplary example, a method for generating a representation of a face is provided. The method includes: obtaining one or more images of a face; generating an encoded expression representing an expression of the face, wherein predetermined characteristics of the face remain constant relative to the encoded expression; mapping the encoded expression to a corresponding expression of the facial model; and generating a representation of the facial model based on the encoded expression.
[0005] As another example, an apparatus for generating a representation of a face is provided. The apparatus includes at least one memory and at least one processor coupled to the at least one memory. The at least one processor is configured to: acquire one or more images of a face; generate an encoded expression representing an expression of the face, wherein predetermined characteristics of the face remain constant relative to the encoded expression; map the encoded expression to a corresponding expression of a facial model; and generate a representation of the facial model based on the encoded expression.
[0006] In another example, a non-transitory computer-readable medium is provided having instructions stored thereon. These instructions, when executed by at least one processor, cause at least one processor to: acquire one or more images of a face; generate an encoded expression representing the facial expression, wherein predetermined characteristics of the face remain constant relative to the encoded expression; map the encoded expression to a corresponding expression of a facial model; and generate a representation of the facial model based on the encoded expression.
[0007] As another example, an apparatus for generating a representation of a face is provided. The apparatus includes: components for acquiring one or more images of a face; components for generating an encoded expression representing an expression of the face, wherein predetermined characteristics of the face remain constant relative to the encoded expression; components for mapping the encoded expression to a corresponding expression of a facial model; and components for generating a representation of the facial model based on the encoded expression.
[0008] In another example, a method for training an expression encoder is provided. The method includes: obtaining a first frame and a second frame, the first frame and the second frame including at least a portion of a face; generating a first expression feature of the first frame, the first expression feature representing a first expression of the face; generating a second expression feature of the second frame, the second expression feature representing a second expression of the face; generating a first viewpoint feature of the first frame, the first viewpoint feature representing a first angle of observation of the face; generating a second viewpoint feature of the second frame, the second viewpoint feature representing a second angle of observation of the face; cross-referencing one of the first expression feature and the second expression feature, or at least one of the first viewpoint feature and the second viewpoint feature; determining a first loss value based on the cross-referencing; and adjusting the feature encoder based on the determined first loss value.
[0009] As another example, an apparatus for training an expression encoder is provided. The apparatus includes: at least one memory; and at least one processor coupled to the at least one memory. The at least one processor is configured to: obtain a first frame and a second frame, the first frame and the second frame including at least a portion of a face; generate a first expression feature of the first frame, the first expression feature representing a first expression of the face; generate a second expression feature of the second frame, the second expression feature representing a second expression of the face; generate a first viewpoint feature of the first frame, the first viewpoint feature representing a first angle of observation of the face; generate a second viewpoint feature of the second frame, the second viewpoint feature representing a second angle of observation of the face; cross-reference one of the first expression feature and the second expression feature, or at least one of the first viewpoint feature and the second viewpoint feature; determine a first loss value based on the cross-reference; and adjust the feature encoder based on the determined first loss value.
[0010] In another example, a non-transitory computer-readable medium having instructions stored thereon is provided. These instructions, when executed by at least one processor, cause at least one processor to: obtain a first frame and a second frame, the first frame and the second frame including at least a portion of a face; generate a first expression feature of the first frame, the first expression feature representing a first expression of the face; generate a second expression feature of the second frame, the second expression feature representing a second expression of the face; generate a first viewpoint feature of the first frame, the first viewpoint feature representing a first angle of observation of the face; generate a second viewpoint feature of the second frame, the second viewpoint feature representing a second angle of observation of the face; cross-reference one of the first expression feature and the second expression feature, or at least one of the first viewpoint feature and the second viewpoint feature; determine a first loss value based on the cross-reference; and adjust a feature encoder based on the determined first loss value.
[0011] As another example, an apparatus for training an expression encoder is provided. The apparatus includes: components for obtaining a first frame and a second frame, the first frame and the second frame including at least a portion of a face; components for generating a first expression feature of the first frame, the first expression feature representing a first expression of the face; components for generating a second expression feature of the second frame, the second expression feature representing a second expression of the face; components for generating a first viewpoint feature of the first frame, the first viewpoint feature representing a first angle of observation of the face; components for generating a second viewpoint feature of the second frame, the second viewpoint feature representing a second angle of observation of the face; components for intersecting one of the first expression feature and the second expression feature, or at least one of the first viewpoint feature and the second viewpoint feature; components for determining a first loss value based on the intersecting; and components for adjusting the feature encoder based on the determined first loss value.
[0012] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to define the scope of the claimed subject matter. This subject matter should be understood with reference to the appropriate portions of the entire specification, any or all drawings, and each claim.
[0013] The foregoing, as well as other features and examples, will become more apparent upon reference to the following description, claims, and drawings. Attached Figure Description
[0014] The following description, with reference to the accompanying drawings, details exemplary examples of this application:
[0015] Figure 1 An architecture for a machine learning model for avatar animation with general pre-trained facial motion encoding, according to various aspects of this disclosure, is illustrated.
[0016] Figure 2An overview of training techniques for machine learning models of avatar animation with general pre-trained facial motion encoding, according to various aspects of this disclosure, is illustrated.
[0017] Figure 3 A technique for encoder training to untangle facial expression information from viewpoint information is illustrated.
[0018] Figure 4 Additional techniques for encoder training to untangle facial expression information from viewpoint information are illustrated according to various aspects of this disclosure.
[0019] Figure 5 Techniques for encoder training to detangle style information from facial expression information according to various aspects of this disclosure are illustrated.
[0020] Figure 6 This is a flowchart illustrating a process for animateting a representation of a face according to various aspects of this disclosure.
[0021] Figure 7 This is a flowchart illustrating a process for animateting a representation of a face according to various aspects of this disclosure.
[0022] Figure 8 It is an exemplary example of a deep learning neural network that can be used by a 3D model training system.
[0023] Figure 9 This is an exemplary example of a convolutional neural network (CNN).
[0024] Figure 10 This is a diagram illustrating an example of a system used to implement certain aspects of this technology. Detailed Implementation
[0025] Certain aspects of this disclosure are provided below. Some of these aspects may be applied independently, and some may be applied in combination, as will be apparent to those skilled in the art. Specific details are set forth in the following description for purposes of explanation in order to provide a thorough understanding of the various aspects of this application. However, it will be apparent that various aspects may be practiced without these specific details. The accompanying drawings and descriptions are not intended to be limiting.
[0026] The following description provides only exemplary aspects and is not intended to limit the scope, applicability, or configuration of this disclosure. Rather, the following description of the exemplary aspects will provide those skilled in the art with a description that can be used to implement the exemplary aspects. It should be understood that various changes may be made to the function and arrangement of the elements without departing from the spirit and scope of this application as set forth in the appended claims.
[0027] The generation of three-dimensional (3D) models of physical objects can be used in many systems and applications, such as extended reality (XR) (e.g., including augmented reality (AR), virtual reality (VR), mixed reality (MR), etc.), robotics, automotive, aviation, 3D scene understanding, object grasping, object tracking, and many other systems and applications. In an AR environment, for example, a user can view an image (also called a frame) that integrates artificial or virtual graphics with the user's natural environment. AR applications allow the manipulation of real-world images to add virtual objects to the image or display virtual objects on a perspective display (making the virtual objects appear to overlay the real-world environment). AR applications can align or register virtual objects with real-world objects (e.g., as observed in an image) in multiple dimensions. For example, real-world objects that exist in reality can be represented using models that are similar to or exactly match the real-world objects. In one example, a model representing a real aircraft located on a runway can be presented through the display of an AR device (e.g., AR glasses, AR head-mounted displays (HMDs), or other devices) while the user continues to view their natural environment through the display. The viewer may be able to manipulate the model while viewing a real-world scene. In another example, models with different colors or physical properties in the AR environment can be used to identify and render actual objects located on a table. In some cases, computer-generated copies of artificial virtual objects that do not exist in reality, or actual objects or structures in the user's natural environment, can also be added to the AR environment.
[0028] The increasing use of facial data in applications (e.g., for XR systems, 3D graphics, security, etc.) has led to a significant demand for systems capable of generating detailed 3D facial models (and 3D models of other objects) efficiently and with high quality. There is also substantial demand for generating 3D models of other types of objects, such as 3D models of vehicles (e.g., for autonomous driving systems), 3D models of room layouts (e.g., for XR applications, for navigation by devices, robots, etc.). Generating detailed 3D models of objects (e.g., 3D facial models) typically requires expensive equipment and the use of multiple cameras in controlled lighting environments, which hinders large-scale data collection.
[0029] Performing 3D facial animation (e.g., animate a 3D model of an object, such as a facial model) can often be challenging, and in many cases, state-of-the-art solutions for facial animation operate based on a training process in which a relatively large number of images are collected for a specific identity (e.g., a specific user) under various lighting conditions to generate a latent code for that trained identity through synthetic analysis. However, such systems may utilize retraining processes for different identities (e.g., different users) or may fail to generalize well across various image types (e.g., for images captured relative to the face in different poses). Instead, one technique is to train a universal encoder to accurately represent facial expressions, regardless of (e.g., depending on) the identity, shading style, and / or viewpoint of the image.
[0030] This document describes systems, apparatuses, electronic devices, methods (also referred to as processes), and computer-readable media (collectively, “Systems and Technologies”) for avatar animation using a general pre-trained facial motion encoding. For example, a machine learning model with a pre-trained visual encoder can be used to extract and encode facial expressions in latent codes (such as vectors) during a pre-training phase. In some cases, the encoded facial expressions can remain constant (e.g., invariant) relative to predetermined characteristics of the face. For example, the encoded facial expressions can remain constant (e.g., face invariant) relative to different viewpoints, color styles, and / or identities. This latent code can then be passed to a fine-tuning phase. This fine-tuning phase can use the encoded expressions to animate a 3D model representation of the user (e.g., an avatar). In some cases, this fine-tuning phase may include a connection layer and a decoder that can be trained to generate and animate avatars with a specific style or appearance. In some cases, audio information may also be used, for example, by the decoder to animate the avatar. For example, audio information can be used to fine-tune a part of the avatar (such as the lips).
[0031] In some cases, the visual encoder may include a motion feature extractor for extracting motion features from the input image. This motion feature extractor may be trained to unwrap facial expression information from identity information, viewpoint information, and / or style information. In some cases, during training, a viewpoint feature extractor may be trained to identify viewpoint features, and this viewpoint feature extractor is used to unwrap viewpoint information (e.g., as provided by the viewpoint features) from facial expression information using a loss value. In some examples, style information may be unwrapped based on a loss value determined by comparing a semantically labeled version of a training image that has been enhanced by altering the color channels of the training image with a version of a training image without color style information.
[0032] Various aspects of this application will be described with reference to the accompanying drawings.
[0033] Figure 1 An architecture for a machine learning model 100 for avatar animation with general pre-trained facial motion encoding, according to various aspects of this disclosure, is illustrated. For example... Figure 1 As shown, the machine learning model 100 can be divided into two general phases: a first phase 102 includes a pre-trained visual encoder 104 for extracting and encoding facial expressions, and a second phase 106 includes fine-tuning for mapping the encoded facial expressions to a user's representation (e.g., an avatar). The visual encoder 104 can be pre-trained to encode facial expressions based on motion, and the visual encoder 104 can be invariant to predetermined characteristics. For example, the visual encoder can be independent of the user's identity, viewpoint (e.g., the angle at which the image of the user is captured), colorization style (e.g., color image, near-infrared image, etc.), etc. (e.g., invariant). For example, the visual encoder 104 can receive images from different viewpoints of the face (e.g., the angle at which the face is viewed or the camera pose at which the image is captured), such as image 108 captured by a camera mounted on a head-mounted device (HMD) 110 or image 112 captured by a non-HMD device (such as a webcam, wireless device, etc.), and for a given expression, the visual encoder 104 can generate the same latent code representing that expression. Additionally, for this expression, the visual encoder 104 can generate the same underlying code representing the expression for multiple different people displaying the expression.
[0034] In some cases, the visual encoder 104 can encode facial expressions as a latent representation of the facial expression (e.g., latent code) as a smooth expression manifold, which can be fine-tuned for any number of downstream applications (such as first application 120 and second application 130). In some cases, the expression manifold can be a distribution of sampled points in a latent space representing the expression, where each sampled point in the latent space represents an aspect of the expression. For example, the expression manifold can be a vector, such as a 256-dimensional vector, where each dimension of the vector represents an aspect of the expression, so changing the entries of the dimension changes the encoded expression. In some cases, the expression manifold can be smooth because it allows transitions between expression states. For example, when a user smiles, their mouth gradually forms a smile over time. The expression manifold can be sensitive enough (e.g., having a sufficient range of dimensions and / or entries of dimensions) to be able to reflect this gradual transition from frame to frame over time. That is, the visual encoder 104 can be able to map the gradual transition from frame to frame to the corresponding latent code (e.g., vector).
[0035] The second phase 106 can be performed by an application (such as a first application 120 and a second application 130), and the second phase 106 may include a set of fully connected layers 122, 132 coupled to decoders 124, 134. In this example, the first application 120 includes a first fully connected layer 122 and a second decoder 124, while the second application 130 includes a second fully connected layer 132 and a second decoder 134. The fully connected layers 122, 132 may be linear models used to map encoded expressions to corresponding expressions of avatars (such as a more realistic first avatar 126 or a more stylized second avatar 136). For example, the fully connected layers 122, 132 may map expressions to morphologies that can be applied to a 3D facial model to generate the expression. Decoders 124, 134 may generate avatars based on the mapped expressions. For example, decoders 124, 134 may obtain a 3D facial model and morph the 3D facial model based on the mapped expressions. In some cases, audio signal 140 may be received, and the received audio signal 140 may be used to enhance (e.g., improve) facial expression generation. For example, audio signal 140 may include the voice of the user represented by the avatar (e.g., emitted by the face). Audio signal 140 may be passed to fully connected layer 142, which may perform affine transformations to help decoders 124, 134 fine-tune the avatar, such as for fine-tuning lip movements. That is, since visual features can be provided by encoded expressions, the movement of the mouth and / or lips can be reconstructed primarily based on the encoded expressions (this helps reconstruct silent mouth movements, such as smiling, frowning, etc.), and audio signal 140 can help reconstruct the fine movements of the mouth and / or lips during speech.
[0036] Figure 2 An overview of training techniques 200 for machine learning models of avatar animations with general pre-trained facial motion encoding, according to various aspects of this disclosure, is illustrated. For example... Figure 2 As shown, technology 200 is based on three types of information: motion information, view information, and style information. In some cases, the first encoder 202 (such as...) Figure 1The encoder 104 can extract motion information of facial expressions in the received image. As discussed above, the first encoder 202 can be a motion feature extractor trained to be identity, viewpoint, and color style / image style invariant, such that the first encoder 202 can extract the same motion (e.g., expression) features from images (e.g., training image 208, near-infrared training image 230) of the same expression from different users captured from different viewpoints (e.g., angle, pose) with different color styles (e.g., red / green / blue (RGB) color image 210, near-infrared image 212, etc.). The extracted motion can be encoded as facial expressions in the image. In some cases, a view feature extractor 206 (e.g., an encoder) can be used to detangle motion information with a specific view (e.g., viewpoint) and / or style.
[0037] During training, the view feature extractor 206 can be used to extract features of the view, and these features can be used to help train the first encoder 202 independently of the angle relative to the face at the time of image capture (e.g., untangling motion with viewpoint information) and the image's coloring style. In some cases, the view feature extractor 206 can be used to train the first encoder 202 via a loss function. In some cases, the view feature extractor 206 may not be used during inference (prediction) time, and the view feature extractor may be discarded after training the first encoder 202. The trained first encoder 202 can be used by any number of downstream applications, for example, to animate an avatar. For example, the same first encoder 202 can be used with any number of downstream applications.
[0038] As discussed above, downstream applications may include a set of fully connected layers 220 and a decoder 222. In some cases, the sets of fully connected layers 220 may be trained by a second encoder 224 to encode image style information for animate creation. Encoding style information may help allow the decoder 222 to add style information, for example, to reproduce training image 208 during training. In some cases, training images with different color styles (such as RGB training image 234 and near-infrared training image 236) may also be input into the second encoder 224 to train the sets of fully connected layers 220. As shown, different sets of fully connected layers 220 can be used in different domains.
[0039] Figure 3An example of a technique used for encoder training 300 to detangle facial expression information from viewpoint information is illustrated. To aid in detanglement of viewpoint and facial expression, it may be useful to use two encoders and cross-combine the extracted features to detangle the facial expression-based features from the viewpoint-based features. In some cases, the loss value can be determined based on the attributes of the cross-combination and a comparison of the extracted features. In some cases, as part of encoder training, the weights of the encoders (e.g., facial expression encoder 304 and / or via angle encoder 306) can be adjusted based on the loss value.
[0040] like Figure 3 As shown, a first training image 302A and a second training image 302B (collectively referred to as training images 302) displaying the same expression from two different angles can each be input into the expression encoder 304 (e.g., Figure 2 The first encoder 202) and the view encoder 306 (e.g., Figure 2 (View feature extractor 206). Expression encoder 304 can generate first expression features of the first training image 302A. Second expression features of image 308A and the second training image 302B 308B. Similarly, the view encoder 306 can generate first view features of the first training image 302A. Second-view features of 310A and the second training image 302B 310B. Since the training images 302 show the same expressions, the expression information can be constrained by the loss function. In some cases, it can be based on the expression loss. To train the facial expression encoder 304, it is necessary to help detangle (e.g., isolate) facial expression information from viewpoint information, so that... And the losses are Loss, which drives the first facial expression feature 308A value and second facial expression features The values for 308B are close because the expressions are the same.
[0041] In some cases, facial crossover loss can also be determined. This helps to further decouple facial expressions from perspective. Facial crossover loss. Based on first-person perspective features 310A and Second-View Features 310B is cross-referenced with 320 and a first facial expression reconstruction image is generated. 312A and second expression reconstruction image 312B is used to obtain, for example, the first-viewpoint features associated with the first training image 302A. 310A can be used with second-view features 310B crosses 320, and uses second-view features. 310B and first facial features 308A to reconstruct the second facial expression image 312B. Since it is assumed that the expression is the same between the first training image 302A and the second training image 302B, the first expression feature is used. 308A and second-view features 310B generates a second facial expression reconstruction image. 312B should reconstruct the second training image 302B. Similarly, for the second training image 302B, the second viewpoint features associated with the second training image 302B... 310B can be used with first-person perspective features 310A crosses 320, and uses first-person perspective features. 310A and second facial expression features 308B (or first facial feature) 308A) to reconstruct the first facial expression image 312A. The image decoder D can be used to determine the difference between the training image and the corresponding facial expression reconstruction image, such that... and Cross-expression loss Then it can be expressed as ,in The loss helps to push the reconstructed expression image to the corresponding training image obtained from different perspectives.
[0042] In some cases, view loss can also be determined. This is to help train the viewpoint encoder 306 to learn how to extract viewpoint features. First-viewpoint features 310A should describe the first pose of the camera used to capture the first training image 302A. Features of 330, and second-view features 310B should describe the second pose of the camera used to capture the second training image 302B. Features of 332. The pose decoder D can be used to determine the difference between the pose expressed in the viewpoint features and the pose of the associated training image, such that... and View loss Then it can be expressed as ,in The loss helps to push the pose expressed in the viewpoint features to the pose of the corresponding training image.
[0043] Figure 4Additional techniques for encoder training 400 to detangle facial expression information from viewpoint information, according to various aspects of this disclosure, are illustrated. To aid in detanglement of viewpoint and facial expression, training using multiple images obtained from the same viewpoint (but including multiple facial expressions) may be useful. For example... Figure 3 As shown, the first training image 402A can be obtained from a specific viewpoint and contains a first expression. The second training image 402B can also be obtained from the same viewpoint, and the second training image 402B may include a second, different expression. The expression encoder 404 can generate a first expression feature representing the first expression in the first training image 402A. 408A, and the expression encoder 404 can generate second expression features representing the second expression in the second training image 402B. 408B. Similarly, the view encoder 406 can generate first view features of the first training image 402A. Second-view features of 410A and the second training image 402B 410B. Since the first training image 402A and the second training image 402B have the same viewpoint, the viewpoint information can be constrained by the loss function. In some cases, it can be based on viewpoint loss. To train the view encoder 306 to help detangle (e.g., isolate) view information from facial expression information, so that And the losses are Loss, which drives first-person perspective features 410A value and second-view feature The values for 410B are close because the viewing angles are the same.
[0044] Additionally, the viewpoint crossover loss can also be determined. This helps to further decouple perspective from facial expression. Perspective crossover loss. Based on the first facial expression features 408A and Second Expression Features 408B performs an interleaving of 420 and generates a first-person reconstructed image. 412A and second expression reconstruction image 412B is used to obtain, for example, the first facial expression feature associated with the first training image 402A. 408A can be used with second facial features 408B crosses 420. Using the second facial expression feature. 408B and first-person perspective features 410A, Second Expression Reconstruction Image Image 412B can be reconstructed for comparison with the second training image 402B. Similarly, the first expression features... 408A and second-view features 410B can be used to reconstruct first facial expression images. 412A is used for comparison with the first training image 402A. The image decoder D can be used to determine the difference between the training image and the corresponding viewpoint reconstructed image, such that... and Cross-expression loss Then it can be expressed as ,in The loss helps to push the reconstructed view image to the corresponding training image containing different expressions.
[0045] In some cases, view loss can also be determined. To aid in training the viewpoint encoder 306. For example, first-view features. 410A should describe the first pose of the camera used to capture the first training image 402A. Features of 430, and second-view features 410B should describe the second pose of the camera used to capture the second training image 402B. Features of 432. Since the viewpoint is the same between the first training image 402A and the second training image 402B, the first viewpoint features... 410A and second-view features 410B should be the same. The pose decoder D can be used to determine the difference between the pose expressed in the viewpoint features and the pose of the associated training image, such that... and View loss Then it can be expressed as ,in The loss helps to push the pose expressed in the viewpoint features to the pose of the corresponding training image.
[0046] Figure 5 Techniques for encoder training 500 to detangle style information from facial expression information according to various aspects of this disclosure are illustrated. In some cases, style information may be information about facial appearance independent of facial expression and head pose. As an example, style information may be information about skin tone, wrinkle appearance, appearance changes due to lighting variations or makeup. To help detangle facial expression and perspective information from color style information (e.g., RGB, grayscale, near IR, etc.), the encoder can be trained by augmenting training images. Augmenting training images can be performed by changing the color channels of training images to generate augmented training images, such as augmented training image 502. During training, augmented training image 502A may be passed to expression encoder 504 (e.g., ...). Figure 2 The first encoder 202) and the view encoder 506 (e.g., Figure 2The view feature extractor 206 generates enhanced facial expression features for the enhanced training image 502A. 508A and enhanced viewing features 510A. In some cases, the alternative expression encoder 512 and the alternative viewpoint encoder 514 can be trained based on a semantically labeled version of the training image 502B that lacks color style information. Examples of semantically labeled versions of the training image may include sketches or segmentation maps of the training image. Sketches of the training image may be obtained based on edge maps extracted from the image or using a sketch-style diffusion model. The alternative expression encoder 512 can generate alternative expression features. 508B, and the alternative view encoder 514 can generate alternative view features. 510B and facial expressions. Since the viewpoint and facial expression information of the semantically labeled versions of the augmented training images 502A and 502B should be identical, style loss can be used as a basis. To train the facial expression encoder 504 and the view encoder 506, where The losses are Loss, which drives the enhancement of facial features 508A and Alternative Facial Features The value of 508B is close, and it promotes enhanced viewpoint features. 510A and Alternative Perspective Features The value of 510B is close.
[0047] In some cases, the decoder 550 (such as for downstream applications) Figure 2 The decoder 222 can be trained to reproduce the latent style of the image. For example, the style encoder 552 can be trained to extract latent style information from the reference style image 556 and used to train the decoder 550 to use enhanced expression features with the reference style from the reference style image 556. 508A and enhanced viewing features 510A reproduces training image 554(I) as reconstructed image 558( Note that decoder 222 can be style-specific. In some cases, the loss used to train the decoder can be the reconstruction loss between the training image 554 and the reconstructed image 558. ), making ( ).
[0048] Figure 6 This is a flowchart illustrating a process 600 for animate a representation of a face according to various aspects of this disclosure. Process 600 may be performed by a computing device (or apparatus) or components of a computing device (e.g., chipset, decoder, etc.). Figure 10The processor 1010, etc., executes the computation. The computing device can be a mobile device (e.g., a mobile phone, etc.), a network-connected wearable device (such as a watch), an extended reality (XR) device (such as a virtual reality (VR) device or an augmented reality (AR) device), a vehicle or a component or system of a vehicle, or other types of computing devices (e.g., Figure 1 HMD 110 Figure 10 The operation of process 600 can be implemented on one or more processors (e.g., computing system 1000). Figure 10 Software components that execute and run on a processor (such as a processor 1010). In some cases, the operation of process 600 can be performed by a processor with... Figure 10 The computing system 1000 is used to implement this.
[0049] At box 602, the computing device (or a component thereof) can obtain one or more images or frames of a face (e.g., Figure 1 Image 108, Image 112).
[0050] At box 604, the computing device (or a component thereof) can generate coded expressions representing facial expressions (e.g., from...). Figure 4 A visual encoder 104). Predetermined features of the face remain constant relative to the encoded expression. In some aspects, the computing device (or a component thereof) may receive an image or frame including at least a portion of a face. The computing device (or a component thereof) may encode motion features of the frame into an encoded expression. The computing device (or a component thereof) may output the encoded expression for transmission. In some cases, the encoded expression may be obtained from a pre-trained visual encoder (e.g., an expression encoder). In some cases, the encoded expression is based on motion features determined from an image of the face. In some examples, the predetermined features of the face include at least one of the face's viewpoint, color style, or identity.
[0051] At box 606, the computing device (or its components) can map the encoded facial expression to the corresponding expression of the facial model. For example, a fully connected layer (such as...) Figure 2 The fully connected layers 122 and 132 can map expressions to deformations that can be applied to a 3D facial model to generate the expression.
[0052] At box 608, a computing device (or a component thereof) can generate a representation of a facial model based on encoded expressions. For example, a decoder (such as...) Figure 1 Decoders 124 and 134 can obtain a 3D facial model and deform the 3D facial model based on mapped expressions. In some cases, this is based on audio signals obtained concurrently with one or more images of the face (e.g., Figure 1 The audio signal 140 is used to enhance the generation of the facial model representation.
[0053] Figure 7 This is a flowchart illustrating a process 700 for animate a representation of a face according to various aspects of this disclosure. Process 700 may be performed by a computing device (or apparatus) or components of a computing device (e.g., chipset, decoder, etc.). Figure 10 The processor 1010, etc., executes the computation. The computing device can be a mobile device (e.g., a mobile phone, etc.), a network-connected wearable device (such as a watch), an extended reality (XR) device (such as a virtual reality (VR) device or an augmented reality (AR) device), a vehicle or a component or system of a vehicle, or other types of computing devices (e.g., Figure 1 HMD 110 Figure 10 The operation of process 700 can be implemented on one or more processors (e.g., computing system 1000). Figure 10 Software components that execute and run on a processor (such as a processor 1010). In some cases, the operation of process 700 can be performed by a processor with... Figure 10 The system is implemented using the architecture of 1000.
[0054] At box 702, the computing device (or a component thereof) can obtain the first frame and the second frame (e.g., Figure 1 Images 108 and 112; Figure 2 Color image 210, near-infrared image 212; Figure 3 Training image 302; Figure 4 Training image 402; Figure 5 The training images (e.g., 554) include at least a portion of the face in the first and second frames.
[0055] At box 704, the computing device (or a component thereof) may generate the first facial expression features of the first frame (e.g., Figure 3 First facial expression features 308A Figure 4 First facial expression features 408A) (for example, via Figure 1 Encoder 104, Figure 2 First encoder 202 Figure 3 304 facial expression encoder Figure 4 The facial expression encoder 404, etc.), the first expression feature represents the first expression of the face.
[0056] At box 706, the computing device (or a component thereof) may generate the second facial expression features of the second frame (e.g., Figure 3 Second facial expression features 308B Figure 4 Second facial expression features (e.g., 408B), the second facial expression feature represents the second facial expression.
[0057] At box 708, the computing device (or a component thereof) may generate first-view features of the first frame (e.g., first-view features). 310A, Figure 4 First-person perspective features 410A, etc. (e.g., via) Figure 2 View feature extractor 206 Figure 3 View encoder 306 Figure 4 (e.g., the view encoder 406), the first view feature represents the first angle at which the face is observed.
[0058] At box 710, the computing device (or a component thereof) may generate second-view features of the second frame (e.g., Figure 3 Second perspective features 310B Figure 4 Second perspective features (410B, etc.), the second perspective feature represents the second angle from which the face is observed.
[0059] At box 712, the computing device (or a component thereof) may cross one of the first facial expression feature and the second facial expression feature, or at least one of the first viewpoint feature and the second viewpoint feature (e.g., Figure 3 Cross 320, Figure 4 (e.g., cross-cutting 420). In some cases, the first expression matches the second expression, and the first viewpoint features cross-cut with the second viewpoint features. In such cases, the computing device (or a component thereof) can generate a first reconstructed image based on the cross-cutting first viewpoint features (e.g., cross-cutting 420, etc.). Figure 3 First expression reconstruction image 312A); Generate a second reconstructed image (e.g., a second expression reconstructed image) based on the cross-referenced second-viewpoint features. 312B); a first loss value is determined based on a comparison between the first reconstructed image and the first frame; and a second loss value (e.g., expression crossover loss) is determined based on a comparison between the second reconstructed image and the second frame. In some examples, the computing device (or a component thereof) may determine a third loss value (e.g., view loss) based on the first and second facial expression features. In some cases, the first angle matches the second angle, and the first expression feature intersects with the second expression feature. In such cases, the computing device (or a component thereof) can generate a first reconstructed image based on the intersected first expression feature (e.g., Figure 4 (412A); a second reconstructed image is generated based on the cross-referenced second expression features (e.g., Figure 4(412B); determining a first loss value based on a comparison between the first reconstructed image and the first frame; and determining a second loss value (e.g., expression crossover loss) based on a comparison between the second reconstructed image and the second frame. In some examples, the computing device (or a component thereof) may determine a third loss value (e.g., view loss) based on first-view features and second-view features. ).
[0060] At box 714, the computing device (or a component thereof) may determine a first loss value (e.g., expression loss) based on the crossover. Cross-expression loss Viewpoint crossover loss Viewpoint loss View loss (etc.). In some cases, a computing device (or a component thereof) can enhance the first frame to generate an enhanced frame (e.g., Figure 5 Enhanced training images 502); Enhanced facial expression features are generated based on the enhanced frames (e.g., Figure 5 Enhanced facial features 508A); based on enhanced frame generation, enhanced viewpoint features (e.g., Figure 5 Enhanced view features 510A); obtain the semantically tagged version of the first frame (e.g., Figure 5 The semantically labeled version of the training image 502B); generating alternative facial features based on the semantically labeled version of the first frame (e.g., Figure 5 Alternative facial features 508B); Generate alternative viewpoint features based on the semantically labeled version of the first frame (e.g., Figure 5 Alternative perspective features 510B); and a third loss value (e.g., style loss) is generated based on the comparison between enhanced facial expression features and alternative facial expression features, and the comparison between enhanced viewpoint features and alternative viewpoint features. In some examples, the computing device (or its components) can enhance the first frame by adjusting its color channels.
[0061] At box 716, the computing device (or a component thereof) may adjust the feature encoder based on the determined first loss value (e.g., by...). Figure 1 Encoder 104, Figure 2 First encoder 202 Figure 2 View feature extractor 206 Figure 3 304 facial expression encoder Figure 3 View encoder 306 Figure 4 404 error in facial expression encoder Figure 4 (e.g., perspective encoder 406).
[0062] Computing devices may include any suitable device, such as mobile devices (e.g., mobile phones), desktop computing devices, tablet computing devices, extended reality (XR) devices or systems (e.g., VR headsets, AR headsets, AR glasses, or other XR devices or systems), wearable devices (e.g., connected watches or smartwatches, or other wearable devices), server computers or systems, computing devices for vehicles or vehicles (e.g., autonomous vehicles), robotic devices, televisions, and / or any other computing device having the resource capability to perform the processes described herein (including process 600). In some cases, a computing device or apparatus may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the processes described herein. In some examples, a computing device may include a display, a network interface configured to communicate and / or receive data, any combination thereof, and / or other components. The network interface may be configured to communicate and / or receive Internet Protocol (IP)-based data or other types of data.
[0063] A component capable of implementing a computing device in a circuit. For example, the component may include electronic circuitry or other electronic hardware, and / or may be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, graphics processing unit (GPU), digital signal processor (DSP), central processing unit (CPU), and / or other suitable electronic circuitry), and / or may include computer software, firmware, or any combination thereof for performing the various operations described herein, and / or may be implemented using computer software, firmware, or any combination thereof for performing the various operations described herein.
[0064] Processes 600 and 700 are illustrated as logic flowcharts, whose operations represent a series of operations that can be implemented by hardware, computer instructions, or a combination thereof. In the context of computer instructions, each operation represents a computer-executable instruction stored on one or more computer-readable storage media, which, when executed by one or more processors, performs the described operation. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc., that perform a specific function or implement a specific data type. The order in which the operations are described is not intended to be construed as limiting, and any number of described operations can be combined in any order and / or in parallel to implement the process.
[0065] Additionally, the processes 600, 700, and / or other processes described herein may be executed under the control of one or more computer systems configured with executable instructions, and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that executes jointly on one or more processors, by hardware, or a combination thereof. As noted above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising multiple instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.
[0066] Figure 8 This is an exemplary example of a deep learning neural network 800 that can be used by a 3D model training system. Input layer 820 includes input data. In one exemplary example, input layer 820 may include data representing pixels of an input video frame. The neural network 800 includes multiple hidden layers 822a, 822b through 822n. Hidden layers 822a, 822b through 822n include "n" hidden layers, where "n" is an integer greater than or equal to one. Multiple hidden layers can be made to include as many layers as needed for a given application. The neural network 800 also includes an output layer 824 that provides the output produced by the processing performed by hidden layers 822a, 822b through 822n. In one exemplary example, output layer 824 may provide a classification of objects in the input video frame. The classification may include a category identifying the type of object (e.g., person, dog, cat, or other object).
[0067] Neural network 800 is a multi-layered neural network composed of interconnected nodes. Each node can represent a piece of information. The information associated with these nodes is shared between different layers, and each layer retains information while processing it. In some cases, neural network 800 may include a feedforward network, in which case there are no feedback connections where the network's output is fed back into itself. In some cases, neural network 800 may include a recurrent neural network, which may have loops that allow information to be carried across nodes as input is read.
[0068] Information can be exchanged between nodes via node-to-node interconnects between layers. Nodes in input layer 820 can activate the node set in the first hidden layer 822a. For example, as shown, each input node in input layer 820 is connected to each node in the first hidden layer 822a. Nodes in hidden layers 822a, 822b, through 822n can transform information by applying an activation function to the information of each input node. The information derived from this transformation can then be passed to nodes in the next hidden layer 822b, activating those nodes, which can then perform their own specified functions. Example functions include convolution, upsampling, data transformation, and / or any other suitable functions. The output of hidden layer 822b can then activate nodes in the next hidden layer, and so on. The output of the last hidden layer 822n can activate one or more nodes in output layer 824, at which the output is provided. In some cases, although a node in neural network 800 (e.g., node 826) is shown as having multiple output lines, the node has a single output and is shown as all lines output from the node representing the same output value.
[0069] In some cases, each node or the interconnection between nodes may have weights, which are a set of parameters derived from the training of the neural network 800. Once the neural network 800 is trained, it can be referred to as a trained neural network, which can be used to classify one or more objects. For example, the interconnection between nodes may represent a piece of information learned about the interconnected nodes. The interconnection may have tunable numerical weights that can be tuned (e.g., based on the training dataset), allowing the neural network 800 to adapt to the input and learn as more and more data is processed.
[0070] The neural network 800 is pre-trained to process features from the data in the input layer 820 using different hidden layers 822a, 822b to 822n in order to provide an output through the output layer 824. In an example where the neural network 800 is used to identify objects in an image, the neural network 800 can be trained using training data that includes both images and labels. For example, training images can be input into the network, where each training image has a label indicating the category of one or more objects in each image (basically, indicating to the network what the objects are and what features they have). In an exemplary example, the training images may include images of the number 2, in which case the label of the image could be [0 0 1 0 0 0 0 0 0 0].
[0071] In some cases, the neural network 800 can use a training process called backpropagation to adjust the weights of its nodes. Backpropagation includes forward pass, loss function, back pass, and weight update. For each training iteration, forward pass, loss function, back pass, and parameter update are performed. This process can be repeated a certain number of iterations for each training image set until the neural network 800 is trained well enough that the weights of each layer are accurately tuned.
[0072] For an example of identifying objects in an image, the forward pass may include passing a training image through a neural network 800. The weights are initially randomized before training the neural network 800. The image may include, for example, a numerical array representing pixels of the image. Each number in the array may include a value from 0 to 255 describing the intensity of the pixel at that location in the array. In one example, the array may include a 28×28×3 numerical array with 28 rows and 28 columns of pixels and 3 color components (such as red, green, and blue, or lightness and two chroma components, etc.).
[0073] For the first training iteration of a neural network 800, the output will likely include values due to the weights being randomly selected during initialization without prioritizing any particular class. For example, if the output is a vector with probabilities that an object includes different classes, the probability values for each class may be equal or at least very similar (e.g., 0.1 for each of ten possible classes). Using the initial weights, the neural network 800 cannot determine low-level features and therefore cannot make an accurate determination of what the object's classification might be. A loss function can be used to analyze the error in the output. Any suitable loss function can be defined. An example of a loss function is Mean Squared Error (MSE). MSE is defined as... It calculates the sum of the actual answer minus half the square of the predicted (output) answer. The loss can be set to equal... The value of .
[0074] For the first training image, the loss (or error) will be high because the actual value will be significantly different from the predicted output. The goal of training is to minimize the loss so that the predicted output matches the training label. The Neural Network 800 performs backpropagation by determining which inputs (weights) contribute most to the network's loss and can adjust the weights to reduce and ultimately minimize the loss.
[0075] The derivative of the loss with respect to the weights (denoted as dL / dW, where W is the weight at a specific layer) can be calculated to determine the weights that contribute the most to the network loss. After calculating the derivative, a weight update can be performed by updating all the weights of the filter. For example, the weights can be updated so that they change in the opposite direction of the gradient. A weight update can be expressed as... η represents the learning rate, where w represents the weights, wi represents the initial weights, and η represents the learning rate. The learning rate can be set to any suitable value, where a high learning rate includes larger weight updates, while a lower value indicates smaller weight updates.
[0076] Neural networks 800 can include any suitable deep network. An example includes a convolutional neural network (CNN), which includes an input layer and an output layer, with multiple hidden layers between the input and output layers. The following section discusses... Figure 8 An example of a CNN is described. The hidden layers of a CNN consist of a series of convolutional layers, non-linear layers, pooling layers (for downsampling), and fully connected layers. Neural networks can include any other deep network besides CNNs, such as autoencoders, deep belief networks (DBNs), recurrent neural networks (RNNs), and so on.
[0077] Figure 9 This is an exemplary example of a Convolutional Neural Network 900 (CNN 900). The input layer 920 of the CNN 900 includes data representing an image. For example, the data could include a numerical array representing the pixels of the image, where each number in the array includes a value from 0 to 255 describing the pixel intensity at that location in the array. Using the previous example above, the array could include a 28×28×3 numerical array with 28 rows and 28 columns of pixels and 3 color components (e.g., red, green, and blue, or lightness and two chroma components, etc.). The image can be passed through a convolutional hidden layer 922a, an optional non-linear activation layer, a pooling hidden layer 922b, and a fully connected hidden layer 922c to obtain an output at the output layer 924. Although... Figure 9 Only one hidden layer from each hidden layer is shown in the diagram, but those skilled in the art will understand that multiple convolutional hidden layers, non-linear layers, pooling hidden layers, and / or fully connected layers may be included in a CNN 900. As previously described, the output may indicate a single category of an object, or may include probabilities that best describe the category of an object in an image.
[0078] The first layer of CNN 900 is a convolutional hidden layer 922a. Convolutional hidden layer 922a analyzes the image data input to layer 920. Each node in convolutional hidden layer 922a is connected to a region of the input image called a receptive field (pixel). Convolutional hidden layer 922a can be thought of as one or more filters (each filter corresponding to a different activation or feature map), where each convolutional iteration of the filter is a node or neuron in convolutional hidden layer 922a. For example, the region of the input image covered by the filter at each convolutional iteration will be the filter's receptive field. In an exemplary example, if the input image consists of a 28×28 array, and each filter (and its corresponding receptive field) is a 5×5 array, then there will be 24×24 nodes in convolutional hidden layer 922a. Each connection between a node and its receptive field learns weights, and in some cases, learns an overall bias, allowing each node to learn to analyze its specific local receptive field in the input image. Each node in hidden layer 922a will have the same weights and biases (referred to as shared weights and shared biases). For example, the filter has a weight (digital) array and the same depth as the input. For the video frame example, the filter would have a depth of 3 (based on the three color components of the input image). An exemplary example of the filter array size is 5×5×3, corresponding to the size of the receptive field of a node.
[0079] The convolutional property of the convolutional hidden layer 922a is due to the fact that each node of the convolutional layer is applied to its corresponding receptive field. For example, the filters of the convolutional hidden layer 922a may begin at the top left corner of the input image array and may convolve around the input image. As noted above, each convolutional iteration of the filters can be considered as a node or neuron of the convolutional hidden layer 922a. In each convolutional iteration, the values of the filters are multiplied by the corresponding number of original pixel values of the image (e.g., a 5×5 filter array is multiplied by a 5×5 array of input pixel values at the top left corner of the input image array). The multiplications from each convolutional iteration can be summed to obtain the sum of that iteration or node. Next, the process continues at the next position in the input image based on the receptive field of the next node in the convolutional hidden layer 922a.
[0080] For example, the filter can move a step size to the next receptive field. The step size can be set to 1 or other suitable amounts. For example, if the step size is set to 1, the filter will move 1 pixel to the right on each convolution iteration. Processing the filter at each unique location in the input volume produces a number representing the filter result at that location, thus determining a sum value for each node of the convolutional hidden layer 922a.
[0081] The mapping from the input layer to the convolutional hidden layer 922a is called an activation map (or feature map). An activation map includes values for each node representing the filter results at each location in the input volume. Activation maps can include arrays containing various sums of values produced by the filter for each iteration of the input volume. For example, if a 5×5 filter is applied to each pixel of a 28×28 input image (with a stride of 1), the activation map would consist of a 24×24 array. The convolutional hidden layer 922a can include several activation maps to identify multiple features in the image. Figure 9 The example shown includes three activation maps. Using these three activation maps, the convolutional hidden layer 922a can detect three different types of features, each of which is detectable across the entire image.
[0082] In some examples, nonlinear hidden layers can be applied after convolutional hidden layers 922a. Nonlinear layers can be used to introduce nonlinearity into a system that has already computed linear operations. An exemplary example of a nonlinear layer is the Corrected Linear Unit (ReLU) layer. A ReLU layer applies the function f(x) = max(0, x) to all values in the input volume, which changes all negative activations to 0. Therefore, ReLU can increase the nonlinearity of CNN 900 without affecting the receptive field of convolutional hidden layers 922a.
[0083] A pooling hidden layer 922b can be applied after the convolutional hidden layer 922a (and, in use, after the non-linear hidden layer). The pooling hidden layer 922b is used to simplify the information in the output of the convolutional hidden layer 922a. For example, the pooling hidden layer 922b can take each activation map output from the convolutional hidden layer 922a and use a pooling function to generate a condensed activation map (or feature map). Max pooling is an example of a function performed by the pooling hidden layer. The pooling hidden layer 922a uses other forms of pooling functions, such as average pooling, L2 norm pooling, or other suitable pooling functions. Pooling functions (e.g., max pooling filters, L2 norm filters, or other suitable pooling filters) are applied to each activation map included in the convolutional hidden layer 922a. Figure 9 In the example shown, three pooling filters are used to convolve the three activation maps in the hidden layer 922a.
[0084] In some examples, max pooling can be used by applying a max pooling filter (e.g., of size 2×2) with a step size (e.g., equal to the size of the filter, such as a step size of 2) to the activation map output from the convolutional hidden layer 922a. The output from the max pooling filter includes the maximum number in each sub-region of the filter convolution. Using a 2×2 filter as an example, each unit in the pooling layer summarizes a region of 2×2 nodes from the previous layer (each node is a value in the activation map). For example, four values (nodes) in the activation map will be analyzed by the 2×2 max pooling filter at each iteration of the filter, with the maximum of the four values being output as the "maximum" value. If such a max pooling filter is applied to an activation filter of size 24×24 nodes from the convolutional hidden layer 922a, the output from the pooling hidden layer 922b will be an array of 12×12 nodes.
[0085] In some examples, L2 norm pooling filters may also be used. L2 norm pooling filters involve calculating the square root of the sum of squares of the values in a 2×2 region (or other suitable region) of the activation map (instead of calculating the maximum value as done in max pooling) and using the calculated value as the output.
[0086] Intuitively, pooling functions (e.g., max pooling, L2-norm pooling, or other pooling functions) determine whether a given feature is found anywhere in a region of an image, discarding the exact location information. This can be done without affecting the results of feature detection, because once a feature has been found, its exact location is less important than its approximate location relative to other features. Max pooling (and other pooling methods) offers the benefit of having far fewer pooling features, thus reducing the number of parameters required in subsequent layers of a CNN 900.
[0087] The final connection in the network is a fully connected layer, which connects each node from the pooling hidden layer 922b to each output node in the output layer 924. Using the example above, the input layer comprises 28×28 nodes encoding the pixel intensity of the input image, the convolutional hidden layer 922a comprises 3×24×24 hidden feature nodes based on applying a 5×5 local receptive field (for filtering) to three activation maps, and the pooling layer 922b comprises 3×12×12 hidden feature nodes based on applying a max-pooling filter to a 2×2 region on each of the three feature maps. Extending this example, the output layer 924 may comprise ten output nodes. In such an example, each node of the 3×12×12 pooling hidden layer 922b is connected to each node of the output layer 924.
[0088] The fully connected layer 922c takes the output of the previous pooling layer 922b (which should represent the activation map of high-level features) and determines the features most relevant to a particular class. For example, the fully connected layer 922c can determine the high-level features most relevant to a particular class and may include weights (nodes) for those high-level features. The product between the weights of the fully connected layer 922c and the pooling hidden layer 922b can be computed to obtain the probabilities for different classes. For example, if CNN 900 is being used to predict whether an object in a video frame is a person, there will be high values in the activation map representing the high-level features of a person (e.g., two legs, a face at the top of the object, two eyes at the upper left and upper right of the face, a nose in the middle of the face, a mouth at the bottom of the face, and / or other features common to people).
[0089] In some examples, the output from output layer 924 may include an M-dimensional vector (M = 10 in the previous example), where M may include the number of categories from which the program must choose when classifying objects in an image. Other example outputs may also be provided. Each number in the N-dimensional vector may represent the probability that an object belongs to a certain category. In an exemplary example, if the 10-dimensional output vector represents objects of ten different categories as [0 0 0.05 0.8 0 0.15 0 0 0 0], then the vector indicates that the probability that the image is an object of the third category (e.g., a dog) is 5%, the probability that the image is an object of the fourth category (e.g., a person) is 80%, and the probability that the image is an object of the sixth category (e.g., a kangaroo) is 15%. The probability of a category can be considered as the confidence level that an object is part of that category.
[0090] Figure 10 This is a diagram illustrating an example of a system used to implement certain aspects of this technology. Specifically, Figure 10 An example of a computing system 1000 is illustrated. This computing system can be any computing device, such as constituting an internal computing system, a remote computing system, a camera, or any component thereof, wherein the components of the system communicate with each other using connection 1005. Connection 1005 can be a physical connection using a bus, or a direct connection to processor 1010, such as in a chipset architecture. Connection 1005 can also be a virtual connection, a networking connection, or a logical connection.
[0091] In some aspects, the computing system 1000 is a distributed system in which the functions described herein can be distributed across a data center, multiple data centers, a peer-to-peer network, etc. In some aspects, one or more of the described system components represent a plurality of such components, each of which performs some or all of the functions described for that component. In some aspects, the components can be physical or virtual devices.
[0092] Example system 1000 includes at least one processing unit (CPU or processor) 1010 and a connection 1005 that couples various system components, including system memory 1015 such as read-only memory (ROM) 1020 and random access memory (RAM) 1025, to processor 1010. Computing system 1000 may include a cache 1012 of high-speed memory that is directly connected to, closely proximate to, or integrated into processor 1010.
[0093] Processor 1010 may include any general-purpose processor and hardware or software services (such as services 1032, 1034, and 1036 stored in storage device 1030 and configured to control processor 1010), as well as dedicated processors in which software instructions are incorporated into the actual processor design. Processor 1010 may be a substantially completely independent computing system containing multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors may be symmetric or asymmetric.
[0094] To enable user interaction, the computing system 1000 includes an input device 1045 that can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice input, etc. The computing system 1000 may also include an output device 1035 that can be one or more of a plurality of output mechanisms. In some instances, a multi-mode system allows a user to provide multiple types of input / output to communicate with the computing system 1000. The computing system 1000 may include a communication interface 1040, which typically governs and manages user input and system output. The communication interface can perform or facilitate the receipt and / or transmission of wired or wireless communications using wired and / or wireless transceivers, including utilizing audio jacks / plugs, microphone jacks / plugs, Universal Serial Bus (USB) ports / plugs, Apple... ® Lightning ® Ports / plugs, Ethernet ports / plugs, fiber optic ports / plugs, dedicated wired ports / plugs, Bluetooth ® Wireless signal transmission, Bluetooth ® Low-power (BLE) wireless signal transmission, IBEACON ®The communication interface 1040 may include one or more Global Navigation Satellite System (GNSS) receivers or transceivers for determining the location of the computing system 1000 based on one or more signals received from one or more satellites associated with one or more GNSS systems. This includes wireless signal transmission, radio frequency identification (RFID) wireless signal transmission, near field communication (NFC) wireless signal transmission, dedicated short range communication (DSRC) wireless signal transmission, 802.11 Wi-Fi wireless signal transmission, wireless local area network (WLAN) signal transmission, visible light communication (VLC) wireless signal transmission, microwave access global interoperability (WiMAX) wireless signal transmission, infrared (IR) wireless signal transmission, public switched telephone network (PSTN) signal transmission, integrated services digital network (ISDN) signal transmission, 3G / 4G / 5G / LTE cellular data network wireless signal transmission, ad hoc network signal transmission, radio wave signal transmission, microwave signal transmission, infrared signal transmission, visible light signal transmission, ultraviolet light signal transmission, wireless signal transmission along the electromagnetic spectrum, or some combination thereof. GNSS systems include, but are not limited to, the U.S. Global Positioning System (GPS), Russia's Global Navigation Satellite System (GLONASS), China's BeiDou Navigation Satellite System (BDS), and Europe's Galileo GNSS. There are no limitations on operation on any particular hardware configuration, and therefore the underlying features here can be easily replaced to obtain improved hardware or firmware configurations as they are developed.
[0095] Storage device 1030 may be a non-volatile and / or non-transitory and / or computer-readable storage device, and may be a hard disk or other type of computer-readable medium capable of storing data accessible by a computer, such as magnetic tape, flash memory cards, solid-state storage devices, digital multifunction disks, cartridges, floppy disks, hard disks, magnetic tapes, magnetic stripes, any other magnetic storage media, flash memory, memristor memory, any other solid-state storage, CD-ROM discs, rewritable CD discs, DVD discs, Blu-ray discs (BDD discs), holographic discs, another optical medium, secure digital (SD) cards, microSD cards, Memory Sticks. ®Cards, smart card chips, EMV chips, Subscriber Identity Module (SIM) cards, mini / micro / nano / micro SIM cards, another integrated circuit (IC) chip / card, random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash EPROM, cache memory (L1 / L2 / L3 / L4 / L5 / L#), resistive random access memory (RRAM / ReRAM), phase change memory (PCM), spin-transfer torque RAM (STT-RAM), another memory chip or cassette and / or combinations thereof.
[0096] Storage device 1030 may include software services, servers, services, etc., which enable the system to perform functions when the code defining such software is executed by processor 1010. In some aspects, hardware services that perform specific functions may include software components for performing functions stored in computer-readable media connected to necessary hardware components such as processor 1010, connection 1005, output device 1035, etc.
[0097] As used herein, the term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. Computer-readable media may include non-transitory media in which data can be stored and which do not include carrier waves and / or transient electronic signals propagating wirelessly or over a wired connection. Examples of non-transitory media may include, but are not limited to, magnetic disks or magnetic tapes, optical storage media (such as compact discs (CDs) or digital versatile discs (DVDs)), flash memory, memory, or memory devices. Computer-readable media may store code and / or machine-executable instructions thereon, which may represent procedures, functions, subroutines, programs, routines, subroutines, modules, software packages, classes, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or hardware circuitry by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc., may be passed, forwarded, or transmitted using any suitable means, including memory sharing, messaging, token passing, network transmission, etc.
[0098] In some respects, computer-readable storage devices, media, and memories may include cables or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer-readable storage media explicitly exclude media such as power consumption, carrier signals, electromagnetic waves, and the signals themselves.
[0099] Specific details are provided in the foregoing description to provide a thorough understanding of the aspects and examples presented herein. However, those skilled in the art will understand that these aspects can be practiced without these specific details. For clarity, in some instances, the technology may be presented as comprising individual functional blocks, including functional blocks containing devices, device components, steps or routines in methods embodied in software or a combination of hardware and software. Additional components may be used in addition to those shown in the figures and / or described herein. For example, circuits, systems, networks, processes and other components may be shown as components in block diagram form to avoid obscuring these aspects in unnecessary detail. In other cases, well-known circuits, processes, algorithms, structures and techniques may be shown without unnecessary detail to avoid obscuring the aspects.
[0100] Various aspects described above can be presented as processes or methods, depicted as flowcharts, diagrams, data flow graphs, structure diagrams, or block diagrams. Although flowcharts can describe operations as sequential processes, many operations within an operation can be executed in parallel or concurrently. Furthermore, the order of operations can be rearranged. A process terminates when its operations are completed, but a process may have additional steps not included in the accompanying diagrams. A process can correspond to a method, function, procedure, subroutine, subroutine, etc. When a process corresponds to a function, its termination may correspond to the function returning to the calling function or the main function.
[0101] The processes and methods described in the examples above can be implemented using stored computer-executable instructions or computer-executable instructions otherwise obtainable from a computer-readable medium. Such instructions may include, for example, instructions and data that configure, or otherwise configure, a general-purpose computer, special-purpose computer, or processing device to perform a function or group of functions. The portion may be accessible via a network of the computer resources used. The computer-executable instructions may be, for example, binary, intermediate format instructions, such as assembly language, firmware, source code, etc. Examples of computer-readable media that can be used to store instructions, information used, and / or information created during the methods according to the described examples include disks or optical discs, flash memory, USB devices with non-volatile memory, networked storage devices, etc.
[0102] Devices implementing the processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented as software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) for performing necessary tasks may be stored in a computer-readable or machine-readable medium. A processor may perform the necessary tasks. Typical examples of form factors include laptop devices, mobile phones (e.g., smartphones or other types of mobile phones), tablet devices or other small form factor personal computers, personal digital assistants, rack-mounted devices, standalone devices, etc. The functionality described herein may also be embodied in peripheral devices or intercalation cards. By further example, such functionality may also be implemented on circuit boards of different chips or different processes executed on a single device.
[0103] Instructions, media for transmitting such instructions, computing resources for executing them, and other structures for supporting such computing resources are example components for providing the functionality described in this disclosure.
[0104] In the foregoing description, aspects of this application have been described with reference to their specific aspects, but those skilled in the art will recognize that this application is not limited thereto. Therefore, although illustrative aspects of this application have been described in detail herein, it is to be understood that various inventive concepts may be embodied and employed in various other ways, and the appended claims are not intended to be construed as including these variations unless limited by prior art. The various features and aspects of the applications described above may be used individually or in combination. Furthermore, aspects may be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of this specification. Therefore, the specification and drawings should be considered illustrative rather than restrictive. For illustrative purposes, the methods are described in a particular order. It should be understood that, in alternative aspects, the methods may be performed in a different order than described.
[0105] Those skilled in the art will understand that, without departing from the scope of this description, the less than (“<”) and greater than (“>”) symbols or terms used herein may be replaced with less than or equal to (“>”) respectively. ") and greater than or equal to (" The symbol ) is used instead.
[0106] When a component is described as being “configured” to perform certain operations, such configuration can be achieved, for example, by designing electronic circuits or other hardware to perform the operations, by programming programmable electronic circuits (e.g., microprocessors or other suitable electronic circuits) to perform the operations, or any combination thereof.
[0107] The phrase “coupled to” means any component that is physically connected directly or indirectly to another component, and / or any component that communicates directly or indirectly with another component (e.g., connected to another component via a wired or wireless connection and / or other suitable communication interface).
[0108] The various exemplary logic blocks, modules, circuits, and algorithm steps described in conjunction with the aspects disclosed herein can be implemented as electronic hardware, computer software, firmware, or combinations thereof. To clearly illustrate this interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and steps have been broadly described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, but such specific implementation decisions should not be construed as departing from the scope of this application.
[0109] The techniques described herein can also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques can be implemented in any of a variety of devices, such as general-purpose computers, wireless communication devices (mobile phones), or integrated circuit devices with multiple uses, including applications in wireless communication devices (mobile phones) and other devices. Any feature described as a module or component can be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, these techniques can be implemented at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium can form part of a computer program product, which may include packaging material. The computer-readable medium may include memory or data storage media, such as random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, etc. Additionally or alternatively, the technology may be implemented at least in part by a computer-readable communication medium that carries or conveys program code in the form of instructions or data structures that can be accessed, read and / or executed by a computer, such as propagated signals or waves.
[0110] The program code can be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such processors can be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; however, in alternatives, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. Therefore, as used herein, the term "processor" may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or means suitable for implementing the techniques described herein.
[0111] Claim language or other languages that state "at least one of" and / or "one or more of" in a set indicate that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, claim language stating "at least one of A and B" or "at least one of A or B" means A, B, or A and B. In another example, claim language stating "at least one of A, B, and C" or "at least one of A, B, or C" means A, B, C, or A and B, or A and C, or B and C, A and B and C, or any repetition is information or data (e.g., A and A, B and B, C and C, A and A and B, etc.), or any other ordering, repetition, or combination of A, B, and C. The language "at least one of the set" and / or "one or more of the set" does not limit the set to the items listed in the set. For example, the language of a claim stating "at least one of A and B" or "at least one of A or B" may refer to A, B, or A and B, and may additionally include items not listed in the set of A and B. The phrases "at least one" and "one or more" are used interchangeably herein.
[0112] Claim language or other languages that state "at least one processor, the at least one processor being configured to," "at least one processor being configured to," "one or more processors, the one or more processors being configured to," etc., indicate that one or more processors (in any combination) are capable of performing associated operations. For example, claim language that states "at least one processor, the at least one processor being configured to: X, Y, and Z" means that a single processor can be used to perform operations X, Y, and Z; or that multiple processors are each assigned a specific subset of tasks of operations X, Y, and Z, such that the multiple processors together perform X, Y, and Z; or that a group of multiple processors work together to perform operations X, Y, and Z. In another example, claim language that states "at least one processor, the at least one processor being configured to: X, Y, and Z" may mean that any single processor can perform only a subset of operations X, Y, and Z.
[0113] When referring to one or more elements that perform functions (e.g., steps of a method), one element may perform all functions, or more than one element may jointly perform these functions. When more than one element jointly performs these functions, each function does not need to be performed by every single element (e.g., different functions may be performed by different elements), and / or each function does not need to be performed by only one element as a whole (e.g., different elements may perform different sub-functions of a function). Similarly, when referring to one or more elements configured to cause another element (e.g., a device) to perform functions, one element may be configured to cause another element to perform all functions, or more than one element may be jointly configured to cause another element to perform these functions.
[0114] When referring to an entity that performs or is configured to perform functions (e.g., steps of a method) (e.g., any entity or device described herein), the entity may be configured to cause one or more elements (individually or collectively) to perform those functions. One or more components of the entity may include at least one memory, at least one processor, at least one communication interface, another component configured to perform one or more of those functions, and / or any combination thereof. When referring to an entity that performs functions, the entity may be configured to cause one component to perform all functions, or to cause more than one component to perform those functions collectively. When the entity is configured to cause more than one component to perform those functions collectively, each function does not need to be performed by every single component (e.g., different functions may be performed by different components), and / or each function does not need to be performed by only one component as a whole (e.g., different components may perform different sub-functions of a function).
[0115] The exemplary aspects of this disclosure include:
[0116] Aspect 1. A method for generating a representation of a face, the method comprising: obtaining one or more images of a face; generating an encoded expression representing an expression of the face, wherein predetermined characteristics of the face remain constant relative to the encoded expression; mapping the encoded expression to a corresponding expression of a facial model; and generating the representation of the facial model based on the encoded expression.
[0117] Aspect 2. The method according to aspect 1, wherein the encoded expression is based on motion features determined from an image of the face.
[0118] Aspect 3. The method according to any one of Aspects 1 to 2, wherein the predetermined characteristics of the face include at least one of the face's perspective, color style, or identity.
[0119] Aspect 4. The method according to any one of Aspects 1 to 4, wherein the generation of the representation of the facial model is enhanced based on audio signals obtained concurrently with the one or more images of the face.
[0120] Aspect 5. The method according to any one of Aspects 1 to 4, the method further comprising: receiving a frame, the frame including at least a portion of a face; encoding motion features of the frame into the encoded expression; and outputting the encoded expression for transmission.
[0121] Aspect 6. A method for training an expression encoder, the method comprising: obtaining a first frame and a second frame, the first frame and the second frame including at least a portion of a face; generating a first expression feature of the first frame, the first expression feature representing a first expression of the face; generating a second expression feature of the second frame, the second expression feature representing a second expression of the face; generating a first viewpoint feature of the first frame, the first viewpoint feature representing a first angle of observation of the face; generating a second viewpoint feature of the second frame, the second viewpoint feature representing a second angle of observation of the face; cross-referencing one of the first expression feature and the second expression feature or at least one of the first viewpoint feature and the second viewpoint feature; determining a first loss value based on the cross-referencing; and adjusting the feature encoder based on the determined first loss value.
[0122] Aspect 7. The method according to aspect 6, wherein the first expression matches the second expression, and wherein the first viewpoint feature intersects with the second viewpoint feature, and the method further includes: generating a first reconstructed image based on the intersected first viewpoint feature; generating a second reconstructed image based on the intersected second viewpoint feature; determining a first loss value based on a comparison between the first reconstructed image and the first frame; and determining a second loss value based on a comparison between the second reconstructed image and the second frame.
[0123] Aspect 8. The method according to aspect 7, the method further comprising determining a third loss value based on the first expression feature and the second expression feature.
[0124] Aspect 9. The method according to any one of Aspects 6 to 8, wherein the first angle matches the second angle, and wherein the first expression feature intersects the second expression feature, and the method further comprises: generating a first reconstructed image based on the intersected first expression feature; generating a second reconstructed image based on the intersected second expression feature; determining a first loss value based on a comparison between the first reconstructed image and the first frame; and determining a second loss value based on a comparison between the second reconstructed image and the second frame.
[0125] Aspect 10. The method according to aspect 9, the method further comprising determining a third loss value based on the first viewpoint features and the second viewpoint features.
[0126] Aspect 11. The method according to any one of Aspects 6 to 10, the method further comprising: enhancing the first frame to generate an enhanced frame; generating enhanced facial expression features based on the enhanced frame; generating enhanced viewpoint features based on the enhanced frame; obtaining a semantically labeled version of the first frame; generating alternative facial expression features based on the semantically labeled version of the first frame; generating alternative viewpoint features based on the semantically labeled version of the first frame; and generating a third loss value based on a comparison between the enhanced facial expression features and the alternative facial expression features, and a comparison between the enhanced viewpoint features and the alternative viewpoint features.
[0127] Aspect 12. The method according to aspect 11, wherein enhancing the first frame includes adjusting the color channels of the first frame.
[0128] Aspect 13. An apparatus for generating a representation of a face, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory, the at least one processor being configured to: acquire one or more images of a face; generate an encoded expression representing an expression of the face, wherein predetermined characteristics of the face remain constant relative to the encoded expression; map the encoded expression to a corresponding expression of a facial model; and generate the representation of the facial model based on the encoded expression.
[0129] Aspect 14. The apparatus according to aspect 13, wherein the encoded expression is based on motion features determined from an image of the face.
[0130] Aspect 15. The apparatus according to any one of Aspects 13 to 14, wherein the predetermined characteristics of the face include at least one of the face's viewing angle, color style, or identity.
[0131] Aspect 16. The apparatus according to any one of aspects 13 to 15, wherein the generation of the representation of the facial model is enhanced based on audio signals obtained concurrently with the one or more images of the face.
[0132] Aspect 17. The apparatus according to any one of Aspects 13 to 16, wherein the processor is further configured to: receive a frame, the frame including at least a portion of a face; encode motion features of the frame into the encoded expression; and output the encoded expression for transmission.
[0133] Aspect 18. An apparatus for training an expression encoder, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory, the at least one processor being configured to: obtain a first frame and a second frame, the first frame and the second frame including at least a portion of a face; generate a first expression feature of the first frame, the first expression feature representing a first expression of the face; generate a second expression feature of the second frame, the second expression feature representing a second expression of the face; generate a first viewpoint feature of the first frame, the first viewpoint feature representing a first angle of observation of the face; generate a second viewpoint feature of the second frame, the second viewpoint feature representing a second angle of observation of the face; cross-reference one of the first expression feature and the second expression feature or at least one of the first viewpoint feature and the second viewpoint feature; determine a first loss value based on the cross-reference; and adjust the feature encoder based on the determined first loss value.
[0134] Aspect 19. The apparatus according to aspect 18, wherein the first expression matches the second expression, and wherein the first viewpoint feature intersects with the second viewpoint feature, and the apparatus further comprises: generating a first reconstructed image based on the intersecting first viewpoint feature; generating a second reconstructed image based on the intersecting second viewpoint feature; determining a first loss value based on a comparison between the first reconstructed image and the first frame; and determining a second loss value based on a comparison between the second reconstructed image and the second frame.
[0135] Aspect 20. The apparatus according to aspect 19, wherein the at least one processor is further configured to determine a third loss value based on the first facial expression feature and the second facial expression feature.
[0136] Aspect 21. The apparatus according to any one of Aspects 18 to 20, wherein the first angle matches the second angle, and wherein the first facial expression feature intersects the second facial expression feature, and wherein the at least one processor is further configured to: generate a first reconstructed image based on the intersecting first facial expression feature; generate a second reconstructed image based on the intersecting second facial expression feature; determine a first loss value based on a comparison between the first reconstructed image and the first frame; and determine a second loss value based on a comparison between the second reconstructed image and the second frame.
[0137] Aspect 22. The apparatus according to aspect 21, wherein the at least one processor is further configured to determine a third loss value based on the first viewpoint feature and the second viewpoint feature.
[0138] Aspect 23. The apparatus according to any one of Aspects 18 to 22, wherein the at least one processor is further configured to: enhance the first frame to generate an enhanced frame; generate enhanced facial expression features based on the enhanced frame; generate enhanced viewpoint features based on the enhanced frame; obtain a semantically labeled version of the first frame; generate alternative facial expression features based on the semantically labeled version of the first frame; generate alternative viewpoint features based on the semantically labeled version of the first frame; and generate a third loss value based on a comparison between the enhanced facial expression features and the alternative facial expression features, and a comparison between the enhanced viewpoint features and the alternative viewpoint features.
[0139] Aspect 24. The apparatus according to aspect 23, wherein, in order to enhance the first frame, the at least one processor is configured to adjust the color channels of the first frame.
[0140] Aspect 25. A non-transitory computer-readable medium having instructions stored thereon, the instructions, when executed by at least one processor, causing the at least one processor to: obtain one or more images of a face; obtain one or more images of a face; generate an encoded expression representing an expression of the face, wherein predetermined characteristics of the face remain constant relative to the encoded expression; map the encoded expression to a corresponding expression of a facial model; and generate a representation of the facial model based on the encoded expression.
[0141] Aspect 26. The non-transitory computer-readable medium according to aspect 25, wherein the encoded expression is based on motion features determined from an image of the face.
[0142] Aspect 27. The non-transitory computer-readable medium according to any one of Aspects 25 to 26, wherein the predetermined characteristics of the face include at least one of the face's viewpoint, color style, or identity.
[0143] Aspect 28. A non-transitory computer-readable medium according to any one of aspects 25 to 26, wherein the generation of the representation of the face from the one or more images is enhanced based on audio signals obtained concurrently with the one or more images of the face.
[0144] Aspect 29. A non-transitory computer-readable medium according to any one of Aspects 25 to 26, wherein the instructions further cause the at least one processor to: receive a frame, the frame including at least a portion of a face; encode motion features of the frame into the encoded expression, wherein the encoded expression is face- and viewpoint-invariant; and output the encoded expression for transmission.
[0145] Aspect 30. A non-transitory computer-readable medium having instructions stored thereon, the instructions, when executed by at least one processor, causing the at least one processor to: obtain a first frame and a second frame, the first frame and the second frame including at least a portion of a face; generate a first expression feature of the first frame, the first expression feature representing a first expression of the face; generate a second expression feature of the second frame, the second expression feature representing a second expression of the face; generate a first viewpoint feature of the first frame, the first viewpoint feature representing a first angle of observation of the face; generate a second viewpoint feature of the second frame, the second viewpoint feature representing a second angle of observation of the face; cross one of the first expression feature and the second expression feature or at least one of the first viewpoint feature and the second viewpoint feature; determine a first loss value based on the crossover; and adjust a feature encoder based on the determined first loss value.
[0146] Aspect 31. The non-transitory computer-readable medium according to aspect 30, wherein the first expression matches the second expression, and wherein the first viewpoint feature intersects with the second viewpoint feature, and the non-transitory computer-readable medium further comprises: generating a first reconstructed image based on the intersected first viewpoint feature; generating a second reconstructed image based on the intersected second viewpoint feature; determining a first loss value based on a comparison between the first reconstructed image and the first frame; and determining a second loss value based on a comparison between the second reconstructed image and the second frame.
[0147] Aspect 32. The non-transitory computer-readable medium according to aspect 31, wherein the instructions cause the at least one processor to determine a third loss value based on the first facial expression feature and the second facial expression feature.
[0148] Aspect 33. A non-transitory computer-readable medium according to any one of aspects 30 to 32, wherein the first angle matches the second angle, and wherein the first facial expression feature intersects the second facial expression feature, and wherein the instructions cause the at least one processor to: generate a first reconstructed image based on the intersecting first facial expression feature; generate a second reconstructed image based on the intersecting second facial expression feature; determine a first loss value based on a comparison between the first reconstructed image and the first frame; and determine a second loss value based on a comparison between the second reconstructed image and the second frame.
[0149] Aspect 34. The non-transitory computer-readable medium according to aspect 33, wherein the instructions cause the at least one processor to determine a third loss value based on the first viewpoint feature and the second viewpoint feature.
[0150] Aspect 35. A non-transitory computer-readable medium according to any one of Aspects 30 to 34, wherein the instructions cause the at least one processor to: enhance the first frame to generate an enhanced frame; generate enhanced facial expression features based on the enhanced frame; generate enhanced viewpoint features based on the enhanced frame; obtain a semantically labeled version of the first frame; generate alternative facial expression features based on the semantically labeled version of the first frame; generate alternative viewpoint features based on the semantically labeled version of the first frame; and generate a third loss value based on a comparison between the enhanced facial expression features and the alternative facial expression features, and a comparison between the enhanced viewpoint features and the alternative viewpoint features.
[0151] Aspect 36. The non-transitory computer-readable medium according to aspect 35, wherein, in order to enhance the first frame, the instructions cause the at least one processor to adjust the color channels of the first frame.
[0152] Aspect 37. An apparatus for generating a representation of a face, the apparatus comprising: one or more components for performing operations according to any one of aspects 1 to 5.
[0153] Aspect 38. An apparatus for training an facial expression encoder, the apparatus comprising one or more components for performing operations according to any one of aspects 6 to 12.
[0154] Aspect 39. A method for generating a representation of a face, the method comprising: obtaining an encoded expression representing an expression of the face, wherein predetermined characteristics of the face remain constant relative to the encoded expression; mapping the encoded expression to a corresponding expression of a facial model; and generating the representation of the facial model based on the encoded expression.
[0155] Aspect 40: An apparatus for generating a representation of a face, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory, the at least one processor being configured to: obtain an encoded expression representing an expression of the face, wherein predetermined characteristics of the face remain constant relative to the encoded expression; map the encoded expression to a corresponding expression of a facial model; and generate the representation of the facial model based on the encoded expression.
[0156] Aspect 41: A non-transitory computer-readable medium having instructions stored thereon, the instructions, when executed by at least one processor, causing the at least one processor to: obtain an encoded expression representing an expression of the face, wherein predetermined characteristics of the face remain constant relative to the encoded expression; map the encoded expression to a corresponding expression of a facial model; and generate a representation of the facial model based on the encoded expression.
Claims
1. A method for generating a representation of a face, the method comprising: Obtain one or more images of a face; Generate an coded expression representing the facial expression, wherein predetermined characteristics of the face remain constant relative to the coded expression; The encoded facial expressions are mapped to the corresponding facial expressions on the facial model; as well as The representation of the facial model is generated based on the encoded expression.
2. The method of claim 1, wherein the encoded expression is based on motion features determined from an image of the face.
3. The method of claim 1, wherein the predetermined characteristics of the face include at least one of the face's perspective, color style, or identity.
4. The method of claim 1, wherein the generation of the representation of the facial model is enhanced based on audio signals obtained concurrently with the one or more images of the face.
5. The method according to claim 1, further comprising: Receive a frame, the frame including at least a portion of a face; The motion features of the frame are encoded into the encoded facial expression; as well as Output the encoded emoticon for sending.
6. A method for training an facial expression encoder, the method comprising: Obtain a first frame and a second frame, wherein the first frame and the second frame include at least a portion of the face; Generate a first facial expression feature for the first frame, wherein the first facial expression feature represents the first facial expression; Generate a second facial expression feature for the second frame, the second facial expression feature representing a second facial expression; Generate a first viewpoint feature for the first frame, wherein the first viewpoint feature represents a first angle from which the face is observed; Generate a second viewpoint feature for the second frame, the second viewpoint feature representing a second angle from which the face is observed; Intersect one of the first facial expression feature and the second facial expression feature or at least one of the first perspective feature and the second perspective feature; The first loss value is determined based on the crossover; as well as The feature encoder is adjusted based on the determined first loss value.
7. The method of claim 6, wherein the first expression matches the second expression, and wherein the first viewpoint feature intersects with the second viewpoint feature, and the method further comprises: The first reconstructed image is generated based on the first viewpoint features after the intersection; A second reconstructed image is generated based on the cross-referenced second-view features; The first loss value is determined based on a comparison between the first reconstructed image and the first frame; and The second loss value is determined based on the comparison between the second reconstructed image and the second frame.
8. The method according to claim 7, further comprising determining a third loss value based on the first expression feature and the second expression feature.
9. The method of claim 6, wherein the first angle matches the second angle, and wherein the first expression feature intersects with the second expression feature, and the method further comprises: A first reconstructed image is generated based on the first facial expression features after cross-referencing; A second reconstructed image is generated based on the cross-referenced second facial expression features; The first loss value is determined based on the comparison between the first reconstructed image and the first frame; as well as The second loss value is determined based on the comparison between the second reconstructed image and the second frame.
10. The method according to claim 9, further comprising determining a third loss value based on the first viewpoint features and the second viewpoint features.
11. The method according to claim 6, further comprising: Enhance the first frame to generate an enhanced frame; Enhanced facial expression features are generated based on the enhanced frames; Enhanced viewpoint features are generated based on the enhanced frames; Obtain the semantic tag version of the first frame; Generate alternative facial features based on the semantic tag version of the first frame; Alternative viewpoint features are generated based on the semantic tag version of the first frame; as well as A third loss value is generated based on the comparison between the enhanced facial expression features and the alternative facial expression features, as well as the comparison between the enhanced viewpoint features and the alternative viewpoint features.
12. The method of claim 11, wherein enhancing the first frame includes adjusting the color channels of the first frame.
13. An apparatus for generating a representation of a face, the apparatus comprising: At least one memory; and At least one processor, coupled to the at least one memory, the at least one processor being configured to: Obtain one or more images of a face; Generate an coded expression representing the facial expression, wherein predetermined characteristics of the face remain constant relative to the coded expression; The encoded facial expressions are mapped to the corresponding facial expressions on the facial model; as well as The representation of the facial model is generated based on the encoded expression.
14. The apparatus of claim 13, wherein the encoded expression is based on motion features determined from an image of the face.
15. The apparatus of claim 13, wherein the predetermined characteristics of the face include at least one of the face's viewing angle, color style, or identity.
16. The apparatus of claim 13, wherein the generation of the representation of the facial model is enhanced based on audio signals obtained concurrently with the one or more images of the face.
17. The apparatus of claim 13, wherein the processor is further configured to: Receive a frame, the frame including at least a portion of a face; Encode the motion features of the frame into the encoded expression; and Output the encoded emoticon for sending.
18. An apparatus for training an facial expression encoder, the apparatus comprising: At least one memory; and At least one processor, coupled to the at least one memory, the at least one processor being configured to: Obtain a first frame and a second frame, wherein the first frame and the second frame include at least a portion of the face; Generate a first facial expression feature for the first frame, wherein the first facial expression feature represents the first facial expression; Generate a second facial expression feature for the second frame, the second facial expression feature representing a second facial expression; Generate a first viewpoint feature for the first frame, wherein the first viewpoint feature represents a first angle from which the face is observed; Generate a second viewpoint feature for the second frame, the second viewpoint feature representing a second angle from which the face is observed; Intersect one of the first facial expression feature and the second facial expression feature or at least one of the first perspective feature and the second perspective feature; The first loss value is determined based on the crossover; as well as The feature encoder is adjusted based on the determined first loss value.
19. The apparatus of claim 18, wherein the first expression matches the second expression, and wherein the first viewpoint feature intersects with the second viewpoint feature, and the apparatus further comprises: The first reconstructed image is generated based on the first viewpoint features after the intersection; A second reconstructed image is generated based on the cross-referenced second-view features; The first loss value is determined based on a comparison between the first reconstructed image and the first frame; and The second loss value is determined based on the comparison between the second reconstructed image and the second frame.
20. The apparatus of claim 19, wherein the at least one processor is further configured to determine a third loss value based on the first facial expression feature and the second facial expression feature.
21. The apparatus of claim 18, wherein the first angle matches the second angle, and wherein the first facial expression feature intersects the second facial expression feature, and wherein the at least one processor is further configured to: A first reconstructed image is generated based on the first facial expression features after cross-referencing; A second reconstructed image is generated based on the cross-referenced second facial expression features; The first loss value is determined based on the comparison between the first reconstructed image and the first frame; as well as The second loss value is determined based on the comparison between the second reconstructed image and the second frame.
22. The apparatus of claim 21, wherein the at least one processor is further configured to determine a third loss value based on the first viewpoint feature and the second viewpoint feature.
23. The apparatus of claim 18, wherein the at least one processor is further configured to: Enhance the first frame to generate an enhanced frame; Enhanced facial expression features are generated based on the enhanced frames; Enhanced viewpoint features are generated based on the enhanced frames; Obtain the semantic tag version of the first frame; Generate alternative facial features based on the semantic tag version of the first frame; Alternative viewpoint features are generated based on the semantic tag version of the first frame; as well as A third loss value is generated based on the comparison between the enhanced facial expression features and the alternative facial expression features, as well as the comparison between the enhanced viewpoint features and the alternative viewpoint features.
24. The apparatus of claim 23, wherein, in order to enhance the first frame, the at least one processor is configured to adjust the color channels of the first frame.
25. A non-transitory computer-readable medium having instructions stored thereon, the instructions causing the at least one processor, when executed, to: Obtain one or more images of a face; Generate an coded expression representing the facial expression, wherein predetermined characteristics of the face remain constant relative to the coded expression; The encoded facial expressions are mapped to the corresponding facial expressions on the facial model; as well as A representation of the facial model is generated based on the encoded expression.
26. The non-transitory computer-readable medium of claim 25, wherein the encoded expression is based on motion features determined from an image of the face.
27. The non-transitory computer-readable medium of claim 25, wherein the predetermined characteristics of the face include at least one of the face's viewpoint, color style, or identity.
28. The non-transitory computer-readable medium of claim 25, wherein the generation of the representation of the face from the one or more images is enhanced based on audio signals obtained concurrently with the one or more images of the face.
29. The non-transitory computer-readable medium of claim 25, wherein the instructions further cause the at least one processor to: Receive a frame, the frame including at least a portion of a face; The motion features of the frame are encoded into the encoded expression, wherein the encoded expression is invariant to facial features and viewpoint; and Output the encoded emoticon for sending.
30. A non-transitory computer-readable medium having instructions stored thereon, the instructions causing the at least one processor, when executed, to: Obtain a first frame and a second frame, wherein the first frame and the second frame include at least a portion of the face; Generate a first facial expression feature for the first frame, wherein the first facial expression feature represents the first facial expression; Generate a second facial expression feature for the second frame, the second facial expression feature representing a second facial expression; Generate a first viewpoint feature for the first frame, wherein the first viewpoint feature represents a first angle from which the face is observed; Generate a second viewpoint feature for the second frame, the second viewpoint feature representing a second angle from which the face is observed; Intersect one of the first facial expression feature and the second facial expression feature or at least one of the first perspective feature and the second perspective feature; The first loss value is determined based on the crossover; as well as The feature encoder is adjusted based on the determined first loss value.