Camera re-projection for faces
By using multiple cameras to capture facial features in a virtual reality headset and combining them with a machine learning model to generate synthetic images, the problem of inaccurate facial expression reprojection in virtual reality systems is solved, improving the realism and naturalness of face-to-face interaction.
Patent Information
- Application Number
- CN202180064993.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-09-22
- Filing Date
- 2021-08-16
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2041-08-16
AI Technical Summary
Existing virtual reality systems struggle to accurately capture and reproject users' facial expressions during face-to-face interactions, especially when users are wearing virtual reality headsets, where the inconsistency between the camera's viewpoint and the rendering viewpoint leads to unnatural facial expressions.
By using multiple inward-facing cameras in a virtual reality headset to capture facial features, combining them with a machine learning model to generate synthetic images, and then using a 3D facial model for mapping and reprojection to render a realistic or avatar-like facial representation.
It enables accurate reprojection of users' facial expressions in a virtual reality environment, improving the realism and naturalness of the user interaction experience.
Smart Images

Figure CN116348919B_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to controls and interfaces for user interaction and experience in virtual reality environments. Background Technology
[0002] Virtual reality (VR) is a computer-generated simulation of an environment (e.g., a 3D environment) that a user can interact with in a seemingly real or physically realistic way. A VR system can be a single device or a group of devices that generate this simulation to display to a user, for example, on a VR headset or some other display device. The simulation can include images, sound, haptic feedback, and / or other senses to mimic a real or fictional environment. As VR becomes increasingly important, its useful applications are rapidly expanding. The most common applications of VR involve games or other interactive content, but other applications such as viewing visual media projects (e.g., photos, videos) for entertainment or training purposes are also emerging. The feasibility of using VR to simulate real-world conversations and other user interactions is also being explored. Summary of the Invention
[0003] This document discloses several different ways of rendering and interacting with virtual (or augmented) reality environments. Artificial reality systems can render artificial environments, which may include virtual spaces rendered to be displayed to one or more users. For example, virtual reality environments or augmented reality environments may be rendered. Users can view and interact within this virtual space and the wider virtual environment in any suitable manner. One objective of the disclosed methods is to reproject a user's facial representation within the artificial reality environment. In a particular embodiment, one or more computing systems can provide a method for reprojecting a user's facial representation within the artificial reality environment. First, the one or more computing systems can receive one or more images of a portion of a user's face. These acquired images can be captured by an inside-out camera coupled to an artificial reality system worn by the user. The one or more computing systems can access a three-dimensional (3D) facial model representing the user's face. The one or more computing systems can identify facial features depicted in the one or more images. The one or more computing systems can determine the camera pose of each camera associated with one or more captured images relative to a 3D facial model based on identified facial features and predetermined feature locations. After determining the poses of the one or more cameras, the one or more computing systems can determine the mapping relationship between the one or more captured images and the 3D facial model. To determine this mapping relationship, the one or more computing systems can project images of said portions of the user's face from the determined camera poses onto the 3D facial model. The one or more computing systems can cause the output image of the user's facial representation to be rendered using the 3D facial model and the mapping relationship between the one or more captured images and the 3D facial model. For example, the one or more computing systems can send instructions to another user's artificial reality system to render the user's facial representation in an artificial reality environment.
[0004] Various embodiments of the present invention may include an artificial reality system or a combination thereof. An artificial reality is a form of reality that has been adjusted in some way before being presented to a user. This artificial reality may include, for example, virtual reality (VR), augmented reality (AR), mixed reality (MR), hybrid reality, or some combination and / or derivative thereof. Artificial reality content may include fully generated content or generated content combined with captured content (e.g., real-world photographs). Artificial reality content may include video, audio, haptic feedback, or some combination thereof, any of which may be presented in single-channel or multi-channel (e.g., stereoscopic video providing a three-dimensional effect to the viewer). Furthermore, in some embodiments, the artificial reality may also be associated with applications, products, accessories, services, or some combination thereof, which are used, for example, to create content in the artificial reality and / or for use in the artificial reality (e.g., performing activities in the artificial reality). Artificial reality systems that deliver artificial reality content can be implemented on a variety of platforms, including head-mounted displays (HMDs) connected to a host computer system, stand-alone HMDs, mobile devices or computing systems, or any other hardware platform capable of delivering artificial reality content to one or more viewers.
[0005] According to one aspect of this disclosure, a method is provided, comprising: one or more computing systems: receiving an image of a portion of a first user's face, wherein the image is captured by a camera coupled to an artificial reality head-mounted device worn by the first user; accessing a three-dimensional (3D) facial model representing the first user's face; identifying one or more facial features depicted in the image; determining a camera pose relative to the 3D facial model based on the identified one or more facial features in the image and predetermined feature positions on the 3D facial model; determining a mapping relationship between the image and the 3D facial model by projecting the image of the portion of the first user's face from the camera pose onto the 3D facial model; and rendering an output image representing the first user's face using at least the 3D facial model and the mapping relationship between the image and the 3D facial model.
[0006] In some embodiments, the method may further include: receiving multiple images corresponding to a second portion of a first user's face, wherein the multiple images are captured by multiple cameras coupled to an artificial reality head-mounted device worn by the first user; synthesizing the images using a machine learning model to generate a synthetic image corresponding to the second portion of the first user's face; and determining a second mapping relationship between the synthetic image and the 3D facial model by projecting the synthetic image of the second portion of the first user's face from a predetermined camera pose onto a 3D facial model.
[0007] In some embodiments, the 3D facial model may be a predetermined 3D facial model representing multiple faces of multiple users.
[0008] In some embodiments, a 3D facial model can be generated by deforming a predetermined 3D facial model representing multiple faces of multiple users, based at least on one or more facial features identified in an image.
[0009] In some embodiments, determining the camera pose may include comparing the location of one or more facial features identified in the image with a predetermined feature location.
[0010] In some embodiments, the output image representing the first user's face can be photorealistic.
[0011] In some embodiments, the method may further include: receiving a second image of a second portion of a first user's face, wherein the second image is captured by a second camera coupled to an artificial reality head-mounted device worn by the first user; identifying one or more second facial features depicted in the second image; determining a second camera pose relative to a 3D facial model based on the identified one or more second facial features in the second image and predetermined feature positions on a 3D facial model; and determining a mapping relationship between the second image and the 3D facial model by projecting the second image of the second portion of the first user's face from the second camera pose onto the 3D facial model, wherein the output image further uses at least the 3D facial model and the mapping relationship between the second image and the 3D facial model.
[0012] In some embodiments, the mapping relationship may be a texture image of that portion of the first user's face, and wherein the texture image is mixed with a predetermined texture corresponding to other portions of the first user's face to generate an output image representing the first user's face.
[0013] In some embodiments, the mapping relationship may be a texture image of a portion of a first user's face, and wherein rendering the first user's facial representation includes: sampling a first point on a predetermined texture corresponding to the first user's facial representation to identify a first color associated with the first point; sampling a second point on the texture image corresponding to the first point on the predetermined texture to identify a second color associated with the second point; and mixing the first color of the first point with the second color of the second point to generate a final color associated with the position corresponding to the first point and the second point.
[0014] In some embodiments, rendering an output image representing a first user's face may include: sending a 3D facial model and a mapping between the image and the 3D facial model to a computing system associated with a second user; and rendering the output image based on the second user's viewpoint relative to the first user.
[0015] According to another aspect of this disclosure, one or more computer-readable non-transitory storage media are provided, comprising software that, when executed, is operable to: receive an image of a portion of a first user's face, wherein the image is captured by a camera coupled to an artificial reality head-mounted device worn by the first user; access a three-dimensional (3D) facial model representing the first user's face; identify one or more facial features depicted in the image; determine a camera pose relative to the 3D facial model based on the identified one or more facial features in the image and multiple predetermined feature positions on the 3D facial model; determine a mapping relationship between the image and the 3D facial model by projecting the image of the portion of the first user's face from the camera pose onto the 3D facial model; and render an output image representing the first user's face using at least the 3D facial model and the mapping relationship between the image and the 3D facial model.
[0016] In some embodiments, when executed, the software may also be operable to: receive a plurality of images corresponding to a second portion of a first user’s face, wherein the plurality of images are captured by a plurality of cameras coupled to an artificial reality head-mounted device worn by the first user; synthesize the images using a machine learning model to generate a synthetic image corresponding to the second portion of the first user’s face; and determine a second mapping relationship between the synthetic image and the 3D facial model by projecting the synthetic image of the second portion of the first user’s face from a predetermined camera pose onto a 3D facial model.
[0017] In some embodiments, the 3D facial model may be a predetermined 3D facial model representing multiple faces of multiple users.
[0018] In some embodiments, the software, when executed, may also be operable to deform a predetermined 3D facial model representing multiple faces of multiple users, based at least on one or more facial features identified in an image.
[0019] In some embodiments, determining the camera pose may include comparing the location of one or more facial features identified in the image with a plurality of predetermined feature locations.
[0020] According to another aspect of this disclosure, a system is provided, comprising: one or more processors; and one or more computer-readable non-transitory storage media coupled to one or more of the one or more processors, and including a plurality of instructions that, when executed by one or more of the one or more processors, are operable to cause the system to: receive an image of a portion of a first user's face, wherein the image is captured by a camera coupled to an artificial reality head-mounted device worn by the first user; access a three-dimensional (3D) facial model representing the first user's face; identify one or more facial features depicted in the image; determine a camera pose relative to the 3D facial model based on the identified one or more facial features in the image and a plurality of predetermined feature positions on the 3D facial model; determine a mapping relationship between the image and the 3D facial model by projecting the image of the portion of the first user's face from the camera pose onto the 3D facial model; and render an output image representing the first user's face using at least the 3D facial model and the mapping relationship between the image and the 3D facial model.
[0021] In some embodiments, the plurality of processors may also be operable, when executing instructions, to: receive a plurality of images corresponding to a second portion of a first user’s face, wherein the plurality of images are captured by a plurality of cameras coupled to an artificial reality head-mounted device worn by the first user; synthesize using a machine learning model to generate a synthetic image corresponding to the second portion of the first user’s face; and determine a second mapping relationship between the synthetic image and the 3D facial model by projecting the synthetic image of the second portion of the first user’s face from a predetermined camera pose onto a 3D facial model.
[0022] In some embodiments, the 3D facial model may be a predetermined 3D facial model representing multiple faces of multiple users.
[0023] In some embodiments, the multiple processors may also be operable, when executing instructions, to deform a predetermined 3D facial model representing multiple faces of multiple users, based at least on one or more facial features identified in an image.
[0024] In some embodiments, determining the camera pose may include comparing the location of one or more facial features identified in the image with a plurality of predetermined feature locations.
[0025] The various embodiments disclosed herein are merely examples, and the scope of this disclosure is not limited to these embodiments. Specific embodiments may include, or exclude, all or some of the components, elements, features, functions, operations, or steps of the embodiments disclosed herein. Embodiments according to the invention are specifically disclosed in the appended claims for methods, storage media, systems, and computer program products, wherein any feature mentioned in one claim class (e.g., method) may also be claimed in another claim class (e.g., system). Dependencies or references in the appended claims are chosen merely for formal reasons. However, protection may also be claimed for any subject matter arising from the intentional reference (in particular multiple dependencies) to any plurality of prior claims, thereby disclosing any combination of the plurality of claims and their features, and regardless of the dependency chosen in the appended claims, any combination of the plurality of claims and their features may be claimed. The subject matter for which protection may be claimed includes not only multiple combinations of the plurality of features set forth in the appended plurality of claims, but also any other combination of the plurality of features in the plurality of claims, wherein each feature mentioned in the plurality of claims may be combined with any other feature in the plurality of claims, or with a combination of multiple other features. Furthermore, protection may be claimed in a single claim for any embodiment and feature of the plurality of embodiments and features described or depicted herein, and / or protection may be claimed for any combination of any embodiment and feature of the plurality of embodiments and features described or depicted herein with any embodiment or feature described or depicted herein, or protection may be claimed for any combination of any embodiment and feature of the plurality of embodiments and features described or depicted herein with any feature of the appended claims. Attached Figure Description
[0026] The patent or application document contains at least one color drawing. Upon request and payment of the necessary fees, the Patent Office will provide a copy of the publication of this patent or patent application with one or more color drawings.
[0027] Figure 1 An example artificial reality system is shown.
[0028] Figure 2 Another example of an artificial reality system is shown.
[0029] Figure 3 Several example camera locations related to the user are shown in the artificial reality system.
[0030] Figure 4 Several example images are shown, captured by multiple cameras of the artificial reality system.
[0031] Figure 5A and Figure 5B An example reprojection area from a camera in an artificial reality system is shown.
[0032] Figure 6 Example rendered 2D images of a 3D model viewed from different perspectives are shown.
[0033] Figure 7A and Figure 7B An example process for reprojecting a facial representation onto a 3D model is shown.
[0034] Figure 8 An example computing system in an artificial reality environment is shown.
[0035] Figure 9 An example method for reprojecting a facial representation onto a facial model is shown.
[0036] Figure 10 An example network environment associated with a virtual reality system is shown.
[0037] Figure 11 An example computer system is shown. Detailed Implementation
[0038] As more people adopt artificial reality (AI) systems, more people will begin using them for various reasons. One use case typically includes face-to-face interaction. These face-to-face interactions can take place within augmented reality (AR) environments, virtual reality (VR) environments, and / or a combination of both. For example, (realistic or non-realistic) avatars or visual representations can be used to represent each user in face-to-face interactions (e.g., virtual meetings). An environment can be presented to each user in which they can see other users during face-to-face interactions. However, these current AI environments between users may not capture facial expressions that people are accustomed to seeing in typical face-to-face interactions. In this respect, facial reprojection within a VR environment can improve the user experience when interacting with other users in an AI environment. However, this may not be a simple problem to solve when the user is wearing a VR headset that partially obscures their face and the image of the user's face is captured from an extreme viewpoint. In this respect, facial models can be used in conjunction with machine learning models to improve camera reprojection of faces within a VR environment.
[0039] In a particular embodiment, the artificial reality system may have one or more cameras capturing the facial features of a user. By way of example, and not limitation, a virtual reality headset may have multiple inside-out cameras capturing the facial features of a user. These inside-out cameras may be used to capture the user's facial features (e.g., a portion of the user's mouth, a portion of the eyes, etc.). Landmarks in the images may be used to deform the facial model to customize it for the user, and the images may be used to create textures for corresponding portions of the facial model. By way of example, and not limitation, an average facial model representing a human face may exist. When a user wears the headset, the headset's cameras may capture images of the user's mouth region. Landmarks in the captured mouth region may be detected and matched with the facial model to determine the pose (position and orientation) of the cameras relative to the facial model. The captured images may be reprojected from the cameras onto the facial model to determine the mapping between the images and the geometry of the facial model (e.g., the images may be used as textures for the facial model). By using facial models, static textures of the user's entire face, and dynamic textures generated based on captured images, artificial reality systems can render images of avatars or realistic representations from the desired viewpoint to present the user's face in a virtual reality environment.
[0040] In certain embodiments, machine learning models can be used to synthesize images of a portion of a user's face. While some cameras in an AIVR headset can clearly capture multiple parts of the face, others may not be able to capture that part clearly and accurately. This is especially true for facial features with complex geometric details (e.g., eyes), and because the camera's viewpoint may differ significantly from the desired rendering viewpoint. Using the cameras of the AIVR headset, images can be captured by each of these cameras and fed into a machine learning model to generate images representing facial portions. By way of example and not limitation, one camera may capture the user's eyes from one angle, while another camera may capture the user's eyes from a different angle. These individual images can be combined to generate a synthetic image that appears to be seen from the front of the eyes. Ground truth images taken from the desired frontal viewpoint can be used to train the machine learning model. The synthetic image can be reprojected from the camera pose used to capture the ground truth image onto a mesh or three-dimensional (3D) facial model (e.g., if the ground truth image was captured by a camera positioned 6 inches directly in front of the user's face and centered between the eyes, the synthetic image will be reprojected from such a camera, rather than from the actual camera used to capture the image).
[0041] In certain embodiments, one or more computing systems may perform the processes described herein. These one or more computing systems may be implemented as social networking systems, third-party systems, artificial reality systems, another computing system, and / or combinations of these computing systems. The one or more computing systems may be coupled to multiple artificial reality systems. These multiple artificial reality systems may be implemented as augmented reality headsets, virtual reality headsets, or mixed reality headsets, etc. In certain embodiments, the one or more computing systems may receive input data from the multiple artificial reality systems. In certain embodiments, the input data may include images captured by one or more cameras coupled to the artificial reality systems.
[0042] In a particular embodiment, the one or more computing systems can receive an image of a portion of a user's face. In a particular embodiment, the one or more computing systems can receive an image of a portion of a user's face from an artificial reality system. This image may be captured by a camera coupled to the artificial reality system. By way of example, and not limitation, the image may be captured by an inside-out camera in an artificial reality headset worn by the user, wherein the image may depict a portion of the user's mouth. Based on the camera pose, the image may correspond to other parts of the user's face. In a particular embodiment, the one or more computing systems can receive multiple images from multiple cameras, the multiple images corresponding to one or more portions of the user's face. By way of example, and not limitation, the one or more computing systems may receive two images corresponding to the user's mouth and one image corresponding to the user's right eye from several inside-out cameras coupled to an artificial reality headset. In a particular embodiment, multiple cameras may be coupled to the artificial reality system relative to the artificial reality system in various camera poses. By way of example, and not limitation, one camera may be placed on the top portion of the artificial reality system, while another camera may be placed on the bottom portion of the artificial reality system. Although this disclosure describes receiving an image of a portion of a user's face in a particular manner, this disclosure contemplates receiving an image of a portion of a user's face in any suitable manner.
[0043] In a particular embodiment, the one or more computing systems may access a three-dimensional (3D) facial model representing a user's face. In a particular embodiment, the one or more computing systems may retrieve the 3D facial model from storage or request it from another computing system. In a particular embodiment, the one or more computing systems may select a 3D facial model based on the user. By way of example, and not limitation, the one or more computing systems may identify user characteristics, such as requesting user input and selecting the 3D facial model best representing the user based on that input. For example, if a user inputs that he / she is six feet tall, African American, and slim, the one or more computing systems may then retrieve the 3D facial model that most accurately represents the user based on that user input. The one or more computing systems may use other factors to access a suitable 3D facial model representing a user's face. In a particular embodiment, the 3D facial model may be a predetermined 3D facial model representing multiple faces of multiple users. By way of example, and not limitation, a generic 3D facial model may be used for all users, for men only, or for women only. In a particular embodiment, the 3D facial model may represent the 3D space that the user's head will occupy in an artificial reality environment. By way of example and not limitation, a 3D facial model can be a mesh, and textures can be applied to that mesh to represent a user's face. Although this disclosure describes accessing a 3D facial model representing a user's face in a particular manner, this disclosure contemplates accessing a 3D facial model representing a user's face in any suitable manner.
[0044] In a particular embodiment, the one or more computing systems can identify one or more facial features depicted in an image. In a particular embodiment, the one or more computing systems can perform a facial feature detection process to identify facial features depicted in an image. By way of example, and not limitation, the one or more computing systems can use a machine learning model to identify cheeks and noses depicted in an image. In a particular embodiment, a specific camera in the artificial reality system can be designated to capture only certain features of the user's face. By way of example, and not limitation, an inside-out camera coupled to the lower part of a virtual reality headset will only capture facial features located below the user's face; therefore, this particular inside-out camera may only identify the mouth or chin. For example, a particular inside-out camera may not be able to identify the eyes. This can reduce the number of features that a particular inside-out camera attempts to identify based on its position on the virtual reality headset. In a particular embodiment, the one or more computing systems can deform a predetermined 3D facial model representing multiple faces of multiple users, based at least on one or more facial features in the identified image. As an example, and not a limitation, the one or more computing systems may, based on identified facial features, determine that a user's face is slightly narrower than a predetermined 3D facial model representing multiple faces of multiple users, and may accordingly deform the predetermined 3D facial model such that it represents a facial model suitable for the user. In a particular embodiment, the one or more computing systems may request an image of the user's face. In a particular embodiment, the one or more computing systems may identify the user's privacy settings to determine whether the one or more computing systems can access the image of the user's face. As an example, and not a limitation, if the user grants permission to the one or more computing systems, the one or more computing systems may access photos associated with the user through an online social network to retrieve an image representing the user's face. One or more retrieved images may be analyzed to determine the static texture representing the user's face. As an example, and not a limitation, one or more retrieved images may be analyzed to determine the location of facial features typically on the user's face. The analyzed image may be compared to a 3D facial model, and the 3D facial model may be deformed based on the analyzed image. Although this disclosure describes identifying one or more facial features depicted in an image in a particular manner, this disclosure contemplates identifying one or more facial features depicted in an image in any suitable manner.
[0045] In a particular embodiment, the one or more computing systems can determine the camera pose relative to a 3D facial model. In a particular embodiment, the one or more computing systems can use one or more facial features identified in the acquired image and predetermined feature locations on the 3D facial model to determine the camera pose. In a particular embodiment, the one or more computing systems can compare the identified locations of facial features depicted in the image with predetermined feature locations. By way of example, and not limitation, the one or more computing systems can identify the location of a user's chin and mouth. The identified locations of the chin and mouth relative to the camera and relative to each other can be used for comparison with multiple predetermined feature locations of the 3D facial model. For example, for a given camera pose, the chin and mouth may be located at specific locations in the acquired image. Taking into account the identified locations of these facial features in the acquired image, the one or more computing systems can determine the position of the camera pose based on comparisons of these identified locations with multiple predetermined feature locations. By way of example and not limitation, if the acquired image contains: 30 pixels from the lower part of the acquired image and 50 pixels from the left side of the acquired image for the chin, and 60 pixels from the upper part of the acquired image and 40 pixels from the right side of the acquired image for the mouth, then the one or more computing systems can subsequently determine that the camera capturing the image may be in a specific camera pose relative to the user's face and the 3D facial model. Considering that a user's head will vary from person to person, comparing the identified facial features with the 3D facial model can allow the one or more computing systems to approximate the camera pose of the camera capturing the image relative to the 3D facial model. Although this disclosure describes determining the camera pose relative to the 3D facial model in a specific manner, this disclosure contemplates determining the camera pose relative to the 3D facial model in any suitable manner.
[0046] In a particular embodiment, the one or more computing systems can determine a mapping between an acquired image and a 3D facial model. In a particular embodiment, the one or more computing systems can determine the mapping between the image and the 3D facial model by projecting an image of a portion of a user's face from a determined camera pose onto the 3D facial model. By way of example and not limitation, the one or more computing systems can acquire a portion of a user's face, such as the user's mouth. Since the camera pose of the camera acquiring the image (e.g., an inside-out camera capturing the image of the user's mouth) is not readily known, the one or more computing systems can determine the camera pose as described herein. Using the camera pose, the one or more computing systems can project the acquired image onto the 3D facial model to determine the mapping. For example, which pixels in the acquired image of the user's mouth belong to a specific location in the 3D facial model. The 3D facial model is used as a mesh to project the image of the user's mouth onto the 3D facial model. In a particular embodiment, the mapping can be a texture image of that portion of the user's face. Although this disclosure describes determining the mapping relationship between acquired images and 3D facial models in a particular manner, this disclosure considers determining the mapping relationship between acquired images and 3D facial models in any suitable manner.
[0047] In a particular embodiment, the computing system can render an output image of a user's facial representation. In a particular embodiment, the computing system can render the output image of the user's facial representation using at least a 3D facial model and a mapping between the acquired image and the 3D facial model. In a particular embodiment, the one or more computing systems can send an instruction to another computing system to render the output image of the user's facial representation. By way of example, and not limitation, the one or more computing systems can send an instruction to a first user's artificial reality system to render an output image of a second user's facial representation. The one or more computing systems can initially receive one or more images of the second user's face acquired from the second user's artificial reality system. As described herein, the one or more computing systems can determine a mapping between the image and a 3D facial model representing the second user's face. The one or more computing systems can send the acquired mapping between the one or more images and the 3D facial model to the first user's artificial reality system. The first user's artificial reality system can render the second user's facial representation based on the received, acquired mapping between the one or more images and the 3D facial model. In a particular embodiment, a rendering package can be sent to the artificial reality system to render the user's facial representation. The rendering package may contain the 3D facial model used, and the mapping between the acquired images and the 3D facial model. As an example, and not a limitation, if the rendering package is used to render a facial representation of a second user, it may include a 3D facial model of the second user and a textured image of the second user's face. Although a general process of reprojecting a user's face has been discussed with respect to a portion of the user's face, the one or more computing systems may receive multiple images corresponding to various parts of the user's face to create a textured image of the user's entire face. As an example, and not a limitation, the one or more computing systems may receive images of the user's eyes, mouth, nose, cheeks, forehead, chin, etc. Each of these images may be used to identify multiple facial features of the user's face and determine a mapping between that image and the user's 3D facial model. This mapping may be used to project the user's entire face onto the 3D facial model. In a particular embodiment, the output image of the user's facial representation may be realistic, making it appear as if the user is looking at another user's face. In a particular embodiment, the output image of the user's facial representation may be an avatar mapped according to the mapping and the 3D facial model. As an example, and not a limitation, mappings can be used to determine the current state of a user's face (e.g., whether the user is smiling, frowning, moving their face in some way, etc.) and reproject that user's face onto an avatar to present a representation of the user's face.In certain embodiments, rendering of the output image may be based on the viewpoint of the user of the artificial reality system, which is rendering the facial representation relative to the user whose face is being rendered. By way of example, and not limitation, the user's facial representation may be considered in the direction in which another user in the artificial reality system is looking at another user. When a user is looking directly at another user, the rendered output image may be that user's facial representation facing that other user. However, if the user is looking at another user indirectly (e.g., from the side), the rendered output image may be a facial representation of that user's side profile. Although this disclosure describes rendering the output image of a user's facial representation in a particular manner, this disclosure contemplates rendering the output image of a user's facial representation in any suitable manner.
[0048] In a particular embodiment, the one or more computing systems can generate a synthetic image corresponding to a portion of a user's face. In a particular embodiment, the one or more computing systems can receive multiple images corresponding to a portion of a user's face. By way of example, and not limitation, the one or more computing systems can receive multiple images corresponding to the user's eyes. In a particular embodiment, the angle at which the camera captures an image of a portion of the user's face can be at an extreme angle relative to the user's face. By way of example, and not limitation, the camera can be positioned close to the user's face. Considering this extreme angle, a direct reprojection of the captured image mapped to a 3D facial model may not be an accurate facial representation and may introduce artifacts during the reprojection process. In this regard, in a particular embodiment, a machine learning model can be used to generate a synthetic image corresponding to one or more portions of the user's face. By way of example, and not limitation, considering that the user's eyes are often obscured by an artificial reality system (e.g., a virtual reality headset), multiple images of the user's eyes at different angles can be aggregated to generate a synthetic image of the user's eyes. In a particular embodiment, the machine learning model can be trained based on ground truth images captured by a camera at a predetermined camera pose relative to the user's face. As an example, and not a limitation, during the training of a machine learning model to synthesize multiple images to generate textures or determine the mapping between synthesized images and a 3D facial model, images of the user's eyes can be captured by an artificial reality system (e.g., by an inside-out camera of the artificial reality system). In a separate process, a camera in a predetermined camera pose can capture images of the user's eyes in an unobstructed manner (e.g., without the user wearing the artificial reality system), which will present a ground truth image for the machine learning model to be compared with the rendered images. The machine learning model can compare a rendering attempt (i.e., the synthesized image) of a portion of the user's face (based on the captured images) with the ground truth image. By training the machine learning model, one or more computational systems can use the machine learning model to synthesize multiple images to generate a synthesized image corresponding to a portion of the user's face (as associated with the multiple images). When reprojecting the texture associated with the synthesized image, the texture can be projected onto the 3D facial model at the same camera pose as the camera that captured the ground truth image. By way of example and not limitation, the one or more computing systems may render an output image of a user's facial representation by at least projecting a synthetic image from a predetermined camera pose onto a 3D facial model. Although this disclosure describes generating a synthetic image corresponding to a portion of a user's face in a particular manner, this disclosure contemplates generating a synthetic image corresponding to a portion of a user's face in any suitable manner.
[0049] In a particular embodiment, the one or more computing systems may blend a texture image with a predetermined texture. In a particular embodiment, the texture image may be a mapping between an image of a portion of a user's face and a 3D facial model of the user. The texture image may be represented by an image and projected onto the 3D facial model to present the portion of the user's face corresponding to the texture image. As an example, and not a limitation, if the captured image is an image of the right side of the user's mouth, the texture image may present the right side of the user's mouth and project it onto the 3D facial model. The projected texture image will be part of a process of reprojecting the user's entire face. For example, the one or more computing systems may receive images of the left side of the user's mouth, the right cheek, the left cheek, etc., to determine the mapping between these captured images and the 3D facial model, and project each of these captured images onto the 3D facial model. In a particular embodiment, the predetermined texture may be based on an image accessed by the one or more computing systems based on the user's privacy settings. The predetermined texture may represent a static texture of the user's face. By way of example, and not limitation, the one or more computing systems can retrieve one or more images of a user's face and generate and / or determine a static texture of the user's face. The predetermined texture can represent what the user's face typically looks like, such as the location of facial features. In a particular embodiment, the one or more computing systems can blend the texture image with the predetermined texture. By way of example, and not limitation, given a predetermined texture, the one or more computing systems can project the texture image and the predetermined texture onto a 3D facial model to accurately reflect the user's face at the current time. For example, if the user is currently smiling, the texture image can correspond to the user's mouth and eyes and can be blended with the predetermined texture of the user's face. The blended texture image with the predetermined image can be used to render a facial representation of a user smiling in an artificial reality environment. In a particular embodiment, during the rendering process, the one or more computing systems can sample a point on the predetermined texture corresponding to the user's facial representation to identify a first color associated with that point. By way of example, and not limitation, the one or more computing systems can sample a pixel corresponding to the user's nose to identify the color of the pixel corresponding to the nose. In a particular embodiment, the one or more computing systems may sample another point on the texture image that corresponds to the point on the predetermined texture to identify another color associated with that point. By way of example, and not limitation, the one or more computing systems may sample the same point on the texture image that corresponds to the predetermined texture. For example, if a point corresponding to a pixel of the nose is sampled on the predetermined texture, then points corresponding to the same pixel of the nose are sampled on the texture image. The color of the sample corresponding to the predetermined texture and the color of the sample corresponding to the texture image may be different.In a particular embodiment, the one or more computing systems may mix a color corresponding to a sample from a predetermined texture and a color corresponding to a sample from a texture image to generate a final color associated with the location corresponding to the sample. By way of example and not limitation, if the sample from the predetermined texture is brown and the sample from the texture image is dark brown, the one or more computing systems may mix these colors and generate a final color between brown and dark brown.
[0050] Figure 1 An example artificial reality system 100 is illustrated. In a particular embodiment, the artificial reality system 100 may include a head-mounted device 104, a controller 106, and a computing system 108. A user 102 may wear the head-mounted device 104, which may display visual artificial reality content to the user 102. The head-mounted device 104 may include an audio device that may provide audio artificial reality content to the user 102. By way of example and not limitation, the head-mounted device 104 may display visual and audio artificial reality content corresponding to a virtual meeting. The head-mounted device 104 may include one or more cameras that may capture images and video of the environment. The head-mounted device 104 may include multiple sensors to determine the head posture of the user 102. The head-mounted device 104 may include a microphone to receive audio input from the user 102. The head-mounted device 104 may be referred to as a head-mounted display (HMD). The controller 106 may include a touchpad and one or more buttons. Controller 106 can receive input from user 102 and forward this input to computing system 108. Controller 106 can also provide haptic feedback to user 102. Computing system 108 can be connected to head-mounted device 104 and controller 106 via cable or wireless connection. Computing system 108 can control head-mounted device 104 and controller 106 to provide artificial reality content to user 102 and receive input from user 102. Computing system 108 can be a standalone host computer system, an onboard computer system integrated with head-mounted device 104, a mobile device, or any other hardware platform capable of providing artificial reality content to user 102 and receiving input from user 102.
[0051] Figure 2An example artificial reality system 200 is illustrated. The artificial reality system 200 can be worn by a user to display an artificial reality environment to the user. The artificial reality system may include displays 202a and 202b to display content to the user. By way of example, and not limitation, the artificial reality system may generate an artificial reality environment of an office space and render an avatar of the user, which includes a facial representation of the user. In a particular embodiment, the artificial reality system may include multiple cameras 204, 206a, and 206b. By way of example, and not limitation, the artificial reality system 200 may include multiple inside-out cameras coupled to the artificial reality system 200 to capture images of various portions of the user's face. By way of example, and not limitation, camera 204 may capture the top portion of the user's face and a portion of the user's eyes. In a particular embodiment, the artificial reality system 200 may send the images captured from the multiple cameras 204, 206a, and 206b to another computing system. This other computing system may process the captured images to send data (e.g., mappings, 3D facial models, etc.) to the other artificial reality system 200 to render the user's facial representation. Although a specific number of components of the artificial reality system 200 are shown, the artificial reality system 200 may include more or fewer components and / or be in different configurations. By way of example and not limitation, the artificial reality system 200 may include additional cameras, and the configuration of all cameras 204, 206a, and 206b may be varied to accommodate the additional cameras. For example, two additional cameras may be present, and cameras 206a and 206b may be positioned closer to displays 202a and 202b.
[0052] Figure 3An environment 300 is shown representing multiple cameras (e.g., cameras 204, 206a, and 206b of an artificial reality system 200) associated with a user's face. In a particular embodiment, several angles 302a to 302c show the distances between the multiple cameras 304, 306a, and 306b associated with the user's face. In a particular embodiment, cameras 304, 306a, and 306b may be coupled to an artificial reality system (not shown). In a particular embodiment, various angles 302a to 302c show cameras 304, 306a, and 306b at a specific distance from the user's face in a specific camera pose. In a particular embodiment, these angles 302a to 302c depict approximate distances between the cameras 304, 306a, and 306b associated with the user's face. Although cameras 304, 306a, and 306b are shown as configured in a particular manner, cameras 304, 306a, and 306b can be reconfigured in different ways. By way of example, and not limitation, cameras 306a and 306b may be positioned lower and closer to the user's face. By way of another example, and not limitation, additional cameras may be present, and cameras 304, 306a, and 306b may be repositioned to accommodate said additional cameras.
[0053] Figure 4 Several example images are shown, captured by multiple cameras of an artificial reality system. These example images are for illustrative purposes only and not as a limitation. Figure 3 The images are captured by cameras 304, 306a, and 306b. In a particular embodiment, camera 304 may capture an image 402 of a portion of the user's face. In a particular embodiment, image 402 may correspond to the top portion of the user's face. In a particular embodiment, camera 306a may capture an image 404 of a portion of the user's face. In a particular embodiment, image 404 may correspond to the user's right eye. In a particular embodiment, camera 306b may capture an image 406 of a portion of the user's face. In a particular embodiment, image 406 may correspond to the user's left eye. Although images 402, 404, and 406 are shown as corresponding to specific portions of the user's face, images 402, 404, and 406 may correspond to different portions of the user's face based on the respective cameras associated with them. In a particular embodiment, any number of images corresponding to a portion of the user's face may be present. By way of example and not limitation, three cameras positioned on the right side of the user's face may be present to capture three images corresponding to the right side of the user's face.
[0054] Figures 5A to 5BAn example reprojection area from a camera in an artificial reality system is shown. In a particular embodiment, an environment 500 representing multiple cameras (e.g., cameras 204, 206a, and 206b of artificial reality system 200) associated with a user's face is shown. Reference Figure 5A Multiple example reprojection regions from cameras (e.g., camera 204) are shown within an environment 500A representing multiple cameras associated with a user's face. In a particular embodiment, reprojection regions at several angles 502a to 502c are shown. In a particular embodiment, at a given angle 502a, the reprojection region 504 of the camera (e.g., camera 204) may include the area between the user's two cheeks. In a particular embodiment, at a given angle 502b, the reprojection region 506 of the camera (e.g., camera 204) may include an area slightly surrounding the user's nose. In a particular embodiment, at a given angle 502c, the reprojection region 508 of the camera (e.g., camera 204) may include the area from the user's forehead to the tip of the user's nose. In a particular embodiment, reprojection regions 504, 506, and 508 may all correspond to the same reprojection region, except that they are shown at different angles 502a to 502c. In a particular embodiment, the one or more computing systems determine the mapping relationship between the acquired images and the 3D facial model or mesh (e.g., image 402 and...). Figure 5A Following the mapping relationship between the 3D facial model or mesh shown, the one or more computing systems can then enable the artificial reality system to render a facial representation of the user's face by projecting the texture image onto reprojection regions 504, 506, and 508 corresponding to the camera that captured the image. Although example reprojection regions 504, 506, and 508 are shown for a specific camera, the camera can have different reprojection regions. By way of example and not limitation, the camera can have a wider field of view and include a larger reprojection region. By way of another example and not limitation, the reprojection region can be smaller if an additional camera is present.
[0055] refer to Figure 5BExample reprojection regions from multiple cameras (e.g., cameras 206a and 206b) are shown within an environment 500B, which represents the multiple cameras associated with the user's face. In a particular embodiment, reprojection regions at several angles 502a to 502c are shown. In a particular embodiment, at a given angle 502a, reprojection regions 510a and 510b of the multiple cameras (e.g., cameras 206a and 206b) may include regions corresponding to the user's eyes. In a particular embodiment, at a given angle 502b, reprojection regions 512a and 512b of the multiple cameras (e.g., cameras 206a and 206b) may include the region between the user's nose and the user's eyes. In a particular embodiment, at a given angle 502c, reprojection regions 514a and 514b of the multiple cameras (e.g., cameras 206a and 206b) (only reprojection region 514a is shown) may include the region from the top to the bottom of the user's eyes. In a particular embodiment, except that the reprojection regions are shown at different angles 502a to 502c, reprojection regions 510a, 510b, 512a, 512b, 514a, and 514b may all correspond to the same reprojection region. In a particular embodiment, the one or more computing systems determine the mapping relationship between the acquired images and the 3D facial model or mesh (e.g., images 404 and 406 and such...). Figure 5B Following the mapping relationship between the 3D facial model or mesh shown, the one or more computing systems can enable the artificial reality system to render a facial representation of the user's face by projecting a texture image onto reprojection regions 510a, 510b, 512a, 512b, 514a, and 514b corresponding to the camera that acquired the image. Although example reprojection regions 510a, 510b, 512a, 512b, 514a, and 514b are shown for a specific camera, the camera may have different reprojection regions. By way of example and not limitation, the camera may have a wider field of view and include a larger reprojection region. By way of another example and not limitation, the reprojection region may be smaller if an additional camera is present.
[0056] Figure 6Example two-dimensional (2D) images of a 3D model viewed from different perspectives are shown. In a particular embodiment, image 602 may represent a real-world view of a user's face. In a particular embodiment, image 604 may represent an example 2D rendered image based on the face shown in image 602 and using a camera with a narrow field of view. By way of example and not limitation, multiple cameras may capture images of various parts of the user's face. These captured images may be used to determine mapping relationships or texture images as shown in image 604. In a particular embodiment, image 606 may represent an example 2D rendered image based on the face shown in image 602 and using a camera with a wider field of view. In a particular embodiment, image 608 may represent another real-world view of the user's face. In a particular embodiment, image 610 may represent an example 2D rendered image based on the face shown in image 608 and using a camera with a narrow field of view. In a particular embodiment, image 612 may represent an example 2D rendered image based on the face shown in image 608 and using a camera with a wider field of view.
[0057] Figure 7A and Figure 7B An example procedure 700 for reprojecting a facial representation onto a 3D model is shown. (Reference) Figure 7A In a particular embodiment, the process 700 may begin with a texture image 702 or a texture image 704, which is determined as described herein based on acquired images and a 3D facial model 710. In a particular embodiment, texture image 702 or 704 may be blended with a predetermined texture 706. As described herein, the predetermined texture 706 may be generated based on multiple images of the user's face. In a particular embodiment, blending texture image 702 or 704 with predetermined texture 706 may generate a new texture 708. Reference Figure 7B The process 700 can continue by projecting the new texture 708 onto the 3D facial model 710. In a particular embodiment, the result of projecting the new texture 708 onto the 3D facial model 710 may be an image 712 showing a representation of the user's face with the new texture 708 projected onto the 3D facial model 710. In a particular embodiment, the blending between the texture image 702 or 704 and the predetermined texture 706 is apparent. In a particular embodiment, during the reprojection process 700, the artificial reality system rendering the facial representation may blend the texture image 702 or 704 with the predetermined texture 706 in such a way that the stark contrast between the two textures is eliminated.
[0058] Figure 8An example computing system 802 in an artificial reality environment 800 is illustrated. In a particular embodiment, the computing system 802 may be implemented as an augmented reality headset, a virtual reality headset, a server, a social networking system, a third-party system, or a combination thereof. Although shown as a single computing system 802, the computing system 802 may be represented by one or more computing systems as described herein. In a particular embodiment, the computing system 802 may be connected to one or more artificial reality systems (e.g., augmented reality headsets, virtual reality headsets, etc.). In a particular embodiment, the computing system 802 may include an input module 804, a feature recognition module 806, a camera pose determination module 808, a mapping module 810, a compositing module 812, a warping module 814, and a reprojection module 816.
[0059] In a particular embodiment, input module 804 may be connected to one or more artificial reality systems in the artificial reality environment 800 to receive input data. In a particular embodiment, the input data may be implemented as images captured from one or more artificial reality systems. By way of example, and not limitation, the artificial reality system may capture images of multiple parts of a user's face from an inside-out camera and send the captured images to computing system 802 as described herein. Input module 804 may send the captured images to other modules of computing system 802. By way of example, and not limitation, input module 804 may send input data to feature recognition module 806, mapping module 810, and / or synthesis module 812.
[0060] In a particular embodiment, the feature recognition module 806 can identify one or more facial features based on input data received from the input module 804. In a particular embodiment, the feature recognition module 806 can perform a facial feature recognition process. In a particular embodiment, the feature recognition module 806 can use a machine learning model to identify one or more facial features within a captured image. In a particular embodiment, the captured image may be associated with a specific camera. By way of example, and not limitation, the captured image may be associated with a top camera coupled to the artificial reality system (e.g., captured by that top camera). The feature recognition module 806 can use this information to reduce the number of facial features to be recognized. By way of example, and not limitation, if the captured image is from a top camera, the feature recognition module 806 will attempt to identify facial features corresponding to the user's eyes, the user's nose, and other facial features located on the upper half of the user's face. In a particular embodiment, the feature recognition module can identify the location associated with the identified facial features. In a particular embodiment, the feature recognition module 806 can send the results of the feature recognition process to other modules of the computing system 802. As an example and not a limitation, the feature recognition module 806 may send the results to the camera pose determination module 808 and / or the deformation module 814.
[0061] In a particular embodiment, the camera pose determination module 808 can determine the camera pose of a camera that has captured an image associated with a facial feature result identified by the feature recognition module 806. In a particular embodiment, as described herein, the camera pose determination module 808 can access a user's 3D facial model. The camera pose determination module 808 can compare the position of facial features in the captured image with predetermined feature positions in the 3D facial model to determine the camera pose corresponding to the captured image. In a particular embodiment, if the camera pose determination module 808 receives multiple facial feature results corresponding to multiple cameras, it can determine the camera pose of each camera that has captured an image sent to the computing system. After determining one or more camera poses, the camera pose determination module 808 can send the determined one or more camera poses associated with one or more captured images to other modules of the computing system 802. By way of example and not limitation, the camera pose determination module 808 can send the determined one or more camera poses to the mapping module 810.
[0062] In a particular embodiment, the mapping module 810 can determine the mapping relationship between the acquired image and the 3D facial model of the user's face by projecting the acquired image from the input module 804 onto the 3D facial model from a determined camera pose corresponding to the acquired image. In a particular embodiment, the mapping module 810 can project multiple acquired images onto the 3D facial model to determine the mapping relationship between the multiple acquired images and the 3D facial model. In a particular embodiment, the mapping module 810 can send the determined mapping relationship to other modules of the computing system 802. By way of example and not limitation, the mapping module 810 can send the mapping relationship to the reprojection module 816.
[0063] In a particular embodiment, the synthesis module 812 may receive input data from the input module 804 corresponding to multiple images of a portion of the user's face. By way of example, and not limitation, the synthesis module 812 may receive multiple images of the user's eyes. In a particular embodiment, as described herein, the synthesis module 812 may synthesize multiple images into a composite image. The synthesis module 812 may send the composite image to other modules of the computing system 802. By way of example, and not limitation, the synthesis module 812 may send the composite image to the reprojection module 816 and / or the mapping module 810. In a particular embodiment, as described herein, the mapping module 810 may determine the mapping relationship between the composite image and the 3D facial model.
[0064] In a particular embodiment, the deformation module 814 may receive the results of identified facial features from the feature recognition module 806. In a particular embodiment, the deformation module 814 may access a 3D facial model representing a user's face. In a particular embodiment, as described herein, the deformation module 814 may deform the 3D facial model. By way of example and not limitation, the deformation module 814 may use the identified facial features compared with the 3D facial model to determine whether the 3D facial model needs to be deformed. For example, if the deformation module 814 determines that the user's nose is 3 inches from the chin, and the 3D facial model has a distance of 2.7 inches between the nose and the chin, then the deformation module 814 may deform the 3D facial model to change the distance between the nose and chin of the 3D facial model to 3 inches. In a particular embodiment, the deformation module 814 may send the result of the deformed 3D facial model to other modules of the computing system 802. By way of example and not limitation, the deformation module 814 may send the deformed 3D facial model to the mapping module 810 and / or the reprojection module 816. In a particular embodiment, the mapping module 810 may use the deformed 3D facial model to determine the mapping relationship between the acquired image and the deformed 3D facial model (instead of the original 3D facial model).
[0065] In a particular embodiment, the reprojection module 816 can receive a mapping relationship from the mapping module. In a particular embodiment, the reprojection module 816 can be connected to one or more artificial reality systems. In a particular embodiment, the reprojection module 816 can generate instructions that cause the artificial reality system to render an output image of the user's facial representation based on the mapping relationship between the acquired image and a 3D facial model (deformed or undeformed). In a particular embodiment, the reprojection module 816 can generate a reprojection package to send to the artificial reality system, the reprojection package including the mapping relationship, the user's 3D facial model, and instructions for rendering the user's facial representation based on the mapping relationship and the 3D facial model.
[0066] Figure 9 An example method 900 for reprojecting a facial representation onto a facial model is illustrated. Method 900 may begin at step 910, in which one or more computing systems may receive an image of a portion of a first user's face. In a particular embodiment, this image may be captured by a camera coupled to an artificial reality head-mounted device worn by the first user. At step 920, the one or more computing systems may access a three-dimensional (3D) facial model representing the first user's face. At step 930, the one or more computing systems may identify one or more facial features depicted in the image. At step 940, the one or more computing systems may determine a camera pose relative to the 3D facial model based on the identified one or more facial features in the image and predetermined feature locations on the 3D facial model. At step 950, the one or more computing systems may determine a mapping between the image and the 3D facial model by projecting the image of that portion of the first user's face from the camera pose onto the 3D facial model. At step 960, the one or more computing systems may at least use a 3D facial model and the mapping relationship between the image and the 3D facial model to render an output image of the first user's facial representation. Where appropriate, certain embodiments may be repeated. Figure 9 One or more steps in the method. Although this disclosure will Figure 9 Specific steps in the method are described and shown as occurring in a specific order, but this disclosure contemplates... Figure 9 Any suitable steps occurring in any suitable order in the method. Furthermore, although this disclosure describes and illustrates example methods for reprojecting a facial representation onto a facial model (including...) Figure 9 The specific steps in the method are not considered here, but this disclosure contemplates any suitable method for reprojecting a facial representation onto a facial model, including any appropriate steps, and where appropriate, the method may include... Figure 9 All steps, some steps, or may be excluded from the method Figure 9Any step in the method. Furthermore, although this disclosure describes and illustrates the execution of... Figure 9 The method may refer to a specific component, device, or system in a specific step, but this disclosure is contemplated for the execution of... Figure 9 Any suitable combination of any suitable component, device, or system in any suitable step of the method.
[0067] Figure 10 An example network environment 1000 associated with a virtual reality system is shown. Network environment 1000 includes a user 1001 interacting with a client system 1030, a social networking system 1060, and a third-party system 1070, which are interconnected via network 1010. Although... Figure 10 A specific arrangement of user 1001, client system 1030, social networking system 1060, third-party system 1070, and network 1010 is shown, but this disclosure considers any suitable arrangement of user 1001, client system 1030, social networking system 1060, third-party system 1070, and network 1010. By way of example and not limitation, two or more of user 1001, client system 1030, social networking system 1060, and third-party system 1070 may bypass network 1010 and be directly connected to each other. As another example, two or more of client system 1030, social networking system 1060, and third-party system 1070 may be physically or logically located in one place, wholly or partially. Furthermore, although... Figure 10 A specific number of users 1001, client systems 1030, social networking systems 1060, third-party systems 1070, and networks 1010 are shown, but this disclosure contemplates any suitable number of users 1001, client systems 1030, social networking systems 1060, third-party systems 1070, and networks 1010. As an example and not a limitation, network environment 1000 may include multiple users 1001, multiple client systems 1030, multiple social networking systems 1060, multiple third-party systems 1070, and multiple networks 1010.
[0068] This disclosure considers any suitable network 1010. By way of example and not limitation, one or more portions of network 1010 may include an ad hoc network, intranet, extranet, virtual private network (VPN), local area network (LAN), wireless LAN (WLAN), wide area network (WAN), wireless wide area network (WWAN), metropolitan area network (MAN), a portion of the Internet, a portion of the Public Switched Telephone Network (PSTN), a cellular telephone network, or a combination of two or more of these networks. Network 1010 may include one or more networks 1010.
[0069] Multiple links 1050 can connect client system 1030, social networking system 1060, and third-party system 1070 to network 1010 or enable client system 1030, social networking system 1060, and third-party system 1070 to each other. This disclosure contemplates any suitable link 1050. In a particular embodiment, one or more links 1050 include one or more wired (e.g., Digital Subscriber Line (DSL) or DataOver Cable Service Interface Specification (DOCSIS)) links, one or more wireless (e.g., Wi-Fi or Worldwide Interoperability for Microwave Access (WiMAX)) links, or one or more optical (e.g., Synchronous Optical Network (SONET) or Synchronous Digital Hierarchy (SDH)) links. In a particular embodiment, one or more links 1050 each include an ad hoc network, intranet, extranet, VPN, LAN, WLAN, WAN, WWAN, MAN, a portion of the Internet, a portion of the PSTN, a cellular network, a satellite communication network, another link 1050, or a combination of two or more such links 1050. In the entire network environment 1000, the links 1050 need not all be identical. One or more first links may differ from one or more second links in one or more respects.
[0070] In a particular embodiment, client system 1030 may be an electronic device comprising hardware, software, or embedded logic components, or a combination of two or more such components, and capable of performing appropriate functions implemented or supported by client system 1030. By way of example and not limitation, client system 1030 may include a computer system, such as a desktop computer, notebook or laptop computer, netbook, tablet computer, e-book reader, GPS device, camera, personal digital assistant (PDA), handheld electronic device, cellular phone, smartphone, virtual reality headset and controller, other suitable electronic devices, or any suitable combination thereof. This disclosure contemplates any suitable client system 1030. Client system 1030 enables network users at client system 1030 to access network 1010. Client system 1030 enables its users to communicate with other users at other client systems 1030. Client system 1030 can generate virtual reality environments for users to interact with content.
[0071] In a particular embodiment, client system 1030 may include a virtual reality (or augmented reality) headset 1032 and one or more virtual reality input devices 1034 (e.g., virtual reality controllers). A user at client system 1030 may wear the virtual reality headset 1032 and use the one or more virtual reality input devices to interact with the virtual reality environment 1036 generated by the virtual reality headset 1032. Although not shown, client system 1030 may also include a separate processing computer and / or any other components of the virtual reality system. The virtual reality headset 1032 may generate the virtual reality environment 1036, which may include system content 1038 (including but not limited to an operating system), such as software or firmware updates, and the virtual reality environment may also include third-party content 1040, such as content from applications or content dynamically downloaded from the Internet (e.g., web page content). The virtual reality headset 1032 may include one or more sensors 1042 (e.g., accelerometers, gyroscopes, magnetometers) to generate sensor data that tracks the position of the virtual reality headset 1032. The virtual reality headset 1032 may also include an eye tracker for tracking the position of the user's eyes or the user's viewing direction. The client system can use data from the one or more sensors 1042 to determine velocity, orientation, and gravity relative to the virtual reality headset. One or more virtual reality input devices 1034 may include one or more sensors 1044 (e.g., accelerometers, gyroscopes, magnetometers, and touch sensors) to generate sensor data tracking the position of the virtual reality input device 1034 and the position of the user's fingers. The client system 1030 may use outside-in tracking, in which a tracking camera (not shown) is positioned outside the virtual reality headset 1032 and within the line of sight of the virtual reality headset 1032. In outside-in tracking, the tracking camera can track the position of the virtual reality headset 1032 (e.g., by tracking one or more infrared light-emitting diode (LED) markers on the virtual reality headset 1032). Alternatively or additionally, the client system 1030 may utilize inside-out tracking, in which a tracking camera (not shown) may be placed on or inside the virtual reality headset 1032. In inside-out tracking, the tracking camera may capture images of the real world around it and may use the changing perspective of the real world to determine its own position in space.
[0072] Third-party content 1040 may include a web browser and may have one or more add-ons, plugins, or other extensions. A user at client system 1030 may enter a Uniform Resource Locator (URL) or direct the web browser to a specific server (e.g., server 1062 or a server associated with third-party system 1070), and the web browser may generate a Hypertext Transfer Protocol (HTTP) request and send that HTTP request to the server. The server may accept the HTTP request and, in response, send one or more Hypertext Markup Language (HTML) files to client system 1030. Client system 1030 may render a web interface (e.g., a webpage) based on the HTML files from the server for presentation to the user. This disclosure considers any suitable source file. As an example and not a limitation, a web interface may be rendered from an HTML file, an Extensible Hypertext Markup Language (XHTML) file, or an Extensible Markup Language (XML) file, depending on specific needs. Such interfaces may also execute scripts, markup languages, and combinations thereof. Throughout this document, references to a web interface, where appropriate, encompass one or more corresponding source files (which the browser may use to render the web interface), and vice versa.
[0073] In a particular embodiment, the social networking system 1060 may be a network-addressable computing system capable of controlling an online social network. The social networking system 1060 may generate, store, receive, and transmit social networking data, such as user profile data, concept profile data, social graph information, or other suitable data related to the online social network. The social networking system 1060 may be accessed directly by other components in the network environment 1000 or via network 1010. By way of example and not limitation, the client system 1030 may use a web browser in third-party content 1040 or a local application associated with the social networking system 1060 (e.g., a mobile social networking application, a messaging application, another suitable application, or any combination thereof) to access the social networking system 1060 directly or via network 1010. In a particular embodiment, the social networking system 1060 may include one or more servers 1062. Each server 1062 may be a single server or a distributed server spanning multiple computers or multiple data centers. Server 1062 can be diverse, such as, but not limited to, web servers, news servers, mail servers, messaging servers, advertising servers, file servers, application servers, exchange servers, database servers, proxy servers, another server suitable for performing the functions or processes described herein, or any combination thereof. In a particular embodiment, each server 1062 may include hardware, software, or embedded logic components, or combinations of two or more such components, for performing appropriate functions implemented or supported by server 1062. In a particular embodiment, social networking system 1060 may include one or more data storage areas 1064. Data storage areas 1064 can be used to store various types of information. In a particular embodiment, the information stored in data storage areas 1064 may be organized according to a particular data structure. In a particular embodiment, each data storage area 1064 may be a relational database, columnar database, correlation database, or other suitable database. Although this disclosure describes or illustrates specific types of databases, this disclosure contemplates any suitable type of database. Specific embodiments may provide multiple interfaces that enable client system 1030, social network system 1060, or third-party system 1070 to manage, retrieve, modify, add, or delete information stored in data storage area 1064.
[0074] In a particular embodiment, the social network system 1060 may store one or more social graphs in one or more data storage areas 1064. In a particular embodiment, the social graph may include multiple nodes—which may include multiple user nodes (each user node corresponds to a specific user) or multiple concept nodes (each concept node corresponds to a specific concept)—and multiple edges connecting these nodes. The social network system 1060 may provide users of the online social network with the ability to communicate and interact with other users. In a particular embodiment, a user can join an online social network via the social network system 1060 and then add connections (e.g., relationships) to some other users in the social network system 1060 that they wish to connect with. Hereinafter, the term "friend" may refer to any other user with whom a user in the social network system 1060 has already formed a connection, association, or relationship through the social network system 1060.
[0075] In a particular embodiment, the social networking system 1060 may provide users with the ability to take action on various types of items or objects supported by the social networking system 1060. By way of example, and not limitation, these items and objects may include groups or social networks to which the user of the social networking system 1060 may belong, events or calendar entries that the user may be interested in, computer-based applications that the user may use, transactions that allow the user to buy or sell items through the service, interactions with advertisements that the user may perform, or other suitable items or objects. Users may interact with anything that can be represented in the social networking system 1060, or with anything that can be represented by an external system 1070, which is decoupled from the social networking system 1060 and coupled to the social networking system 1060 via network 1010.
[0076] In a particular embodiment, the social networking system 1060 may be able to link various entities. By way of example and not limitation, the social networking system 1060 may enable multiple users to interact with each other and receive content from third-party systems 1070 or other entities, or allow users to interact with these entities through an application programming interface (API) or other communication channels.
[0077] In certain embodiments, third-party system 1070 may include one or more types of servers, one or more data storage areas, one or more interfaces (including but not limited to APIs), one or more web services, one or more content sources, one or more networks, or any other suitable components (e.g., servers may communicate with these components). Third-party system 1070 may be operated by an entity different from the entity operating social networking system 1060. However, in certain embodiments, social networking system 1060 and third-party system 1070 may operate collaboratively to provide social networking services to users of social networking system 1060 or third-party system 1070. In this sense, social networking system 1060 may provide a platform or backbone network that other systems (e.g., third-party system 1070) can use to provide social networking services and functionality to users on the Internet.
[0078] In a particular embodiment, the third-party system 1070 may include a third-party content object provider. The third-party content object provider may include one or more content object sources that can be transmitted to the client system 1030. As an example, and not a limitation, the content object may include information about things or activities that the user is interested in, such as movie showtimes, movie reviews, restaurant reviews, restaurant menus, product information and reviews, or other suitable information. As another example, and not a limitation, the content object may include incentive content objects, such as coupons, discount vouchers, gift certificates, or other suitable incentive content objects.
[0079] In a particular embodiment, the social networking system 1060 also includes user-generated content objects, which can enhance user interaction with the social networking system 1060. User-generated content can include any content that a user can add, upload, send, or "post" to the social networking system 1060. As an example, and not a limitation, a user transmits a post from the client system 1030 to the social networking system 1060. A post can include data such as status updates or other text data, location information, photos, videos, links, music, or other similar data or media. Content can also be added to the social networking system 1060 by a third party via a "communication channel" such as a news feed or stream.
[0080] In a particular embodiment, the social networking system 1060 may include various servers, subsystems, programs, modules, logs, and data storage areas. In a particular embodiment, the social networking system 1060 may include one or more of the following: a web server, an action logger, an API request server, a relevance and ranking engine, a content object classifier, a notification controller, action logs, third-party content object exposure logs, an inference module, an authorization / privacy server, a search module, an ad targeting module, a user interface module, a user profile storage area, a contact storage area, a third-party content storage area, or a location storage area. The social networking system 1060 may also include suitable components, such as a web interface, security mechanisms, load balancers, failover servers, management and network operations consoles, other suitable components, or any suitable combination thereof. In a particular embodiment, the social networking system 1060 may include one or more user profile storage areas for storing user profiles. User profiles may include, for example, biometric information, demographic information, behavioral information, social information, or other types of descriptive information (e.g., work experience, educational history, hobbies or preferences, interests, close relationships, or location). Interest information may include interests associated with one or more categories. These categories can be generic or specific. As an example, and not a limitation, if a user “likes” an article about a shoe brand, the category could be that brand, or the generic categories “shoes” or “clothing.” A contact store can be used to store contact information about users. Contact information can indicate users who have similar or shared work experience, group memberships, hobbies, educational history, or are in any way related to or share common attributes. Contact information can also include user-defined connections between different users and content (both internal and external). A web server can be used to link the social networking system 1060 to one or more client systems 1030 or one or more third-party systems 1070 via network 1010. The web server can include a mail server or other messaging functionality for receiving and sending messages between the social networking system 1060 and one or more client systems 1030. An API request server can allow third-party systems 1070 to access information from the social networking system 1060 by calling one or more APIs. An action recorder can be used to receive communications from the web server regarding user actions on or outside the social networking system 1060. By combining action logs, logs of third-party content objects exposed to users can be maintained. The notification controller can provide information about content objects to the client system 1030. Information can be pushed to the client system 1030 as a notification, or retrieved from the client system 1030 in response to a received request.An authorization server can be used to enforce one or more privacy settings for users of the social networking system 1060. A user's privacy settings determine how specific information associated with that user can be shared. The authorization server can allow users, for example, by setting appropriate privacy settings, to choose whether or not their actions are recorded by the social networking system 1060 or shared with other systems (e.g., third-party system 1070). A third-party content object storage area can be used to store content objects received from third parties (such as third-party system 1070). A location storage area can be used to store location information received from the client system 1030 associated with the user. An advertising targeting module can combine social information, current time, location information, or other suitable information to deliver relevant advertisements to users in the form of notifications.
[0081] Figure 11 An example computer system 1100 is illustrated. In a particular embodiment, one or more computer systems 1100 perform one or more steps of one or more methods described or illustrated herein. In a particular embodiment, one or more computer systems 1100 provide the functionality described or illustrated herein. In a particular embodiment, software running on one or more computer systems 1100 performs one or more steps of one or more methods described or illustrated herein, or provides the functionality described or illustrated herein. The particular embodiments include one or more portions of one or more computer systems 1100. Throughout this document, references to computer systems may include computing devices and vice versa, where appropriate. Furthermore, references to computer systems may include one or more computer systems, where appropriate.
[0082] This disclosure contemplates any suitable number of computer systems 1100. This disclosure contemplates computer systems 1100 employing any suitable physical form. By way of example and not limitation, computer system 1100 may be an embedded computer system, a system-on-chip (SOC), a single-board computer system (SBC) (e.g., a computer-on-module (COM) or system-on-module (SOM)), a desktop computer system, a laptop or notebook computer system, an interactive self-service machine, a mainframe, a network of computer systems, a mobile phone, a personal digital assistant (PDA), a server, a tablet computer system, or a combination of two or more of these computer systems. Where appropriate, computer system 1100 may include one or more computer systems 1100; computer system 1100 may be single or distributed; spanning multiple locations; spanning multiple machines; spanning multiple data centers; or located in the cloud (which may include one or more cloud components in one or more networks). Where appropriate, one or more computer systems 1100 can perform one or more steps of the methods described or illustrated herein without significant space or time constraints. By way of example and not limitation, one or more computer systems 1100 can perform one or more steps of the methods described or illustrated herein in real time or in batch mode. Where appropriate, one or more computer systems 1100 can perform one or more steps of the methods described or illustrated herein at different times or in different locations.
[0083] In a particular embodiment, computer system 1100 includes a processor 1102, memory 1104, storage device 1106, input / output (I / O) interface 1108, communication interface 1110, and bus 1112. Although this disclosure describes and illustrates a particular computer system having a particular number of components in a particular arrangement, this disclosure contemplates any suitable computer system having any suitable number of components in any suitable arrangement.
[0084] In a particular embodiment, processor 1102 includes hardware for executing a plurality of instructions, such as those that constitute a computer program. By way of example, and not limitation, to execute the plurality of instructions, processor 1102 may retrieve (or read) these instructions from internal registers, internal cache, memory 1104, or storage device 1106; decode and execute these instructions; and then write one or more results to the internal registers, internal cache, memory 1104, or storage device 1106. In a particular embodiment, processor 1102 may include one or more internal caches for data, instructions, or addresses. Where appropriate, this disclosure contemplates processor 1102 including any suitable number of suitable internal caches. By way of example, and not limitation, processor 1102 may include one or more instruction caches, one or more data caches, and one or more translation lookaside buffers (TLBs). The plurality of instructions in the instruction cache may be copies of the plurality of instructions in memory 1104 or storage device 1106, and the instruction cache may accelerate the retrieval of those instructions by processor 1102. The data in the data cache may be a copy of the data in memory 1104 or storage device 1106 for operation by instructions executed at processor 1102; the result of a previous instruction executed at processor 1102 for access by subsequent instructions executed at processor 1102, or for writing to memory 1104 or storage device 1106; or the data in the data cache may be other suitable data. The data cache can accelerate read or write operations of processor 1102. Multiple TLBs can accelerate virtual address translation of processor 1102. In a particular embodiment, processor 1102 may include one or more internal registers for data, instructions, or addresses. Where appropriate, this disclosure contemplates processor 1102 including any suitable number of suitable internal registers. Where appropriate, processor 1102 may include one or more arithmetic logic units (ALUs); processor 1102 may be a multi-core processor, or may include one or more processors 1102. Although this disclosure describes and illustrates specific processors, this disclosure contemplates any suitable processor.
[0085] In a particular embodiment, memory 1104 includes main memory for storing instructions to be executed by processor 1102 or data to be operated by processor 1102. By way of example and not limitation, computer system 1100 may load multiple instructions from storage device 1106 or another source (e.g., another computer system 1100) into memory 1104. Processor 1102 may then load these instructions from memory 1104 into internal registers or internal cache memory. To execute these instructions, processor 1102 may retrieve and decode these instructions from internal registers or internal cache memory. During or after the execution of these instructions, processor 1102 may write one or more results (which may be intermediate or final results) into internal registers or internal cache memory. Processor 1102 may then write one or more of those results into memory 1104. In a particular embodiment, processor 1102 executes only instructions in one or more internal registers or one or more internal caches, or in memory 1104 (different from memory device 1106 or other locations), and operates only on data in one or more internal registers or one or more internal caches, or in memory 1104 (different from memory device 1106 or other locations). One or more memory buses (each memory bus may include an address bus and a data bus) couple processor 1102 to memory 1104. As described below, bus 1112 may include one or more memory buses. In a particular embodiment, one or more memory management units (MMUs) are located between processor 1102 and memory 1104 and facilitate access to memory 1104 requested by processor 1102. In a particular embodiment, memory 1104 includes random access memory (RAM). Where appropriate, the RAM may be volatile memory. Where appropriate, the RAM may be dynamic RAM (DRAM) or static RAM (SRAM). Furthermore, where appropriate, the RAM may be a single-port RAM or a multi-port RAM. This disclosure contemplates any suitable RAM. Where appropriate, memory 1104 may include one or more memories 1104. Although this disclosure describes and illustrates specific memories, this disclosure contemplates any suitable memory.
[0086] In a particular embodiment, storage device 1106 includes a mass storage device for data or instructions. By way of example and not limitation, storage device 1106 may include a hard disk drive (HDD), a floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these storage devices. Where appropriate, storage device 1106 may include removable or non-removable (or fixed) media. Where appropriate, storage device 1106 may be internal or external to computer system 1100. In a particular embodiment, storage device 1106 is a non-volatile solid-state memory. In a particular embodiment, storage device 1106 includes read-only memory (ROM). Where appropriate, the ROM may be a mask-programmable ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM), or flash memory, or a combination of two or more of these ROMs. This disclosure contemplates a high-capacity storage device 1106 in any suitable physical form. Where appropriate, storage device 1106 may include one or more storage control units facilitating communication between processor 1102 and storage device 1106. Where appropriate, storage device 1106 may include one or more storage devices 1106. Although this disclosure describes and illustrates specific storage devices, this disclosure contemplates any suitable storage device.
[0087] In a particular embodiment, I / O interface 1108 includes hardware, software, or both hardware and software that provide one or more interfaces for communication between computer system 1100 and one or more I / O devices. Where appropriate, computer system 1100 may include one or more of these I / O devices. These one or more I / O devices enable communication between a person and computer system 1100. By way of example and not limitation, I / O devices may include a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, still camera, stylus, input pad, touchscreen, trackball, camera, another suitable I / O device, or a combination of two or more of these I / O devices. I / O devices may include one or more sensors. This disclosure contemplates any suitable I / O device and any suitable I / O interface 1108 for such I / O devices. Where appropriate, I / O interface 1108 may include one or more device or software drivers that enable processor 1102 to drive one or more of these I / O devices. Where appropriate, I / O interface 1108 may include one or more I / O interfaces 1108. Although this disclosure describes and illustrates specific I / O interfaces, this disclosure considers any suitable I / O interface.
[0088] In a particular embodiment, the communication interface 1110 includes hardware, software, or both, providing one or more interfaces for communication (e.g., packet-based communication) between the computer system 1100 and one or more other computer systems 1100 or one or more networks. By way of example, and not limitation, the communication interface 1110 may include a network interface controller (NIC) or network adapter for communicating with Ethernet or other wire-based networks, or a wireless NIC (WNIC) or wireless adapter for communicating with wireless networks such as Wi-Fi networks. This disclosure contemplates any suitable network and any suitable communication interface 1110 for that network. By way of example, and not limitation, the computer system 1100 may communicate with one or more portions of an ad hoc network, a personal area network (PAN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), or the Internet, or a combination of two or more of these networks. One or more portions of one or more of these networks may be wired or wireless. As an example, computer system 1100 may communicate with a wireless PAN (WPAN) (e.g., Bluetooth WPAN), a Wi-Fi network, a Wi-Fi Max network, a cellular telephone network (e.g., a Global System for Mobile Communication (GSM) network), or other suitable wireless networks, or a combination of two or more of these networks. Where appropriate, computer system 1100 may include any suitable communication interface 1110 for any of these networks. Where appropriate, communication interface 1110 may include one or more communication interfaces 1110. Although specific communication interfaces are described and illustrated in this disclosure, any suitable communication interface is contemplated in this disclosure.
[0089] In a particular embodiment, bus 1112 includes hardware, software, or both hardware and software that couple multiple components of computer system 1100 to each other. By way of example and not limitation, bus 1112 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a front-side bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infiniband interconnect, a low-pin-count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCIe) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association local (VLB) bus, or another suitable bus, or a combination of two or more of these buses. Where appropriate, bus 1112 may include one or more buses 1112. Although this disclosure describes and illustrates a particular bus, this disclosure considers any suitable bus or interconnect.
[0090] In this document, where appropriate, a computer-readable non-transitory storage medium may include one or more semiconductor-based integrated circuits (ICs) or other integrated circuits (e.g., field-programmable gate arrays (FPGAs) or application-specific ICs (ASICs)), hard disk drives (HDDs), hybrid hard drives (HHDs), optical discs, optical disc drives (ODDs), magneto-optical disks, magneto-optical disk drives, floppy disks, floppy disk drives (FDDs), magnetic tape, solid-state drives (SSDs), RAM drives, secure digital cards, secure digital drives, any other suitable computer-readable non-transitory storage media, or any suitable combination of two or more of these storage media. Where appropriate, a computer-readable non-transitory storage medium may be volatile, non-volatile, or a combination of volatile and non-volatile computer-readable non-transitory storage media.
[0091] In this document, unless otherwise expressly indicated or the context otherwise indicates, “or” is inclusive rather than exclusive. Therefore, in this document, unless otherwise expressly indicated or the context otherwise indicates, “A or B” means “A, B, or both A and B”. Furthermore, unless otherwise expressly indicated or the context otherwise indicates, “and” is both common and separate. Therefore, in this document, unless otherwise expressly indicated or the context otherwise indicates, “A and B” means “A and B, commonly or separately”.
[0092] The scope of this disclosure covers all changes, substitutions, variations, alterations, and modifications to the exemplary embodiments described or illustrated herein that will be understood by those skilled in the art. The scope of this disclosure is not limited to the exemplary embodiments described or illustrated herein. Furthermore, although this disclosure describes and illustrates various embodiments herein as including specific components, elements, features, functions, operations, or steps, those skilled in the art will understand that any embodiment in these embodiments may include any combination or arrangement of any component, element, feature, function, operation, or step described or illustrated anywhere herein. Moreover, references in the appended claims to apparatus or systems, or components in apparatus or systems (that are adapted, arranged, enabled, configured, implemented, operable, or usable to perform a particular function) cover that apparatus, system, or component (whether or not the apparatus, system, component, or the particular function is activated, turned on, or unlocked), provided that the apparatus, system, or component is so adapted, arranged, enabled, configured, implemented, operable, or usable. Furthermore, although this disclosure describes or illustrates specific embodiments to provide particular advantages, specific embodiments may not provide these advantages, or may provide some or all of these advantages.
Claims
1. A method, the method comprising: Consists of one or more computing systems: Receive an image of a portion of a first user's face, wherein the image is captured by a camera coupled to a first artificial reality head-mounted device worn by the first user; Access represents a three-dimensional 3D facial model of the first user's face; Identify one or more facial features depicted in the image; Based on the comparison of the identified one or more facial features in the image with predetermined feature positions on the 3D facial model, the relative pose of the camera of the first artificial reality head-mounted device with respect to the 3D facial model is determined. The mapping relationship between the image of the portion of the first user's face and the 3D facial model is determined by projecting the image of the portion of the first user's face from the relative pose of the camera of the first artificial reality head-mounted device onto the 3D facial model. A reprojection package is generated using a reprojection module. The reprojection package includes the mapping relationship, the 3D facial model of the first user, and instructions for rendering the facial representation of the first user based on the mapping relationship and the 3D facial model. The first AI reality headset sends the reprojection packet to the second AI reality headset; and An output image of the facial representation of the first user is rendered from the viewpoint of the second user using the second artificial reality head-mounted device relative to the first user, wherein the output image is rendered using at least the 3D facial model and the mapping relationship between the image and the 3D facial model.
2. The method according to claim 1, further comprising: Receive multiple images corresponding to a second portion of the face of the first user, wherein the multiple images are captured by multiple cameras coupled to the first artificial reality head-mounted device worn by the first user; Synthesis is performed using a machine learning model to generate a synthetic image corresponding to the second portion of the first user's face; and A second mapping relationship between the synthetic image and the 3D facial model is determined by projecting the synthetic image of the second portion of the first user's face from a predetermined camera pose onto the 3D facial model.
3. The method according to claim 1 or 2, wherein, The 3D facial model is a pre-defined 3D facial model representing multiple faces of multiple users.
4. The method according to claim 1 or 2, wherein, The 3D facial model is generated in the following way: The predetermined 3D facial model representing multiple faces of multiple users is deformed based at least on the one or more facial features identified in the image.
5. The method according to claim 1 or 2, wherein, Determining the relative pose of the camera includes comparing the position of one or more facial features identified in the image with the position of the predetermined feature.
6. The method according to claim 1 or 2, wherein, The output image representing the face of the first user has a realistic feel.
7. The method according to claim 1 or 2, further comprising: Receive a second image of a second portion of the face of the first user, wherein the image is captured by a second camera coupled to the first artificial reality head-mounted device worn by the first user; Identify one or more second facial features depicted in the second image; Based on the identified one or more second facial features in the second image and the predetermined feature positions on the 3D facial model, a second camera pose relative to the 3D facial model is determined; and The mapping relationship between the second image and the 3D facial model is determined by projecting a second image of the second portion of the first user's face from the pose of the second camera onto the 3D facial model, wherein the output image also uses at least the 3D facial model and the mapping relationship between the second image and the 3D facial model.
8. The method according to claim 1 or 2, wherein, The mapping relationship is a texture image of a portion of the face of the first user, and wherein: i. The texture image is mixed with a predetermined texture corresponding to other parts of the first user's face to generate the output image representing the first user's face; and / or ii. Rendering the facial representation of the first user includes: A first point on a predetermined texture corresponding to the facial representation of the first user is sampled to identify a first color associated with the first point; Sampling is performed on a second point on the texture image corresponding to the first point on the predetermined texture to identify a second color associated with the second point; and The first color at the first point is mixed with the second color at the second point to generate a final color associated with the position corresponding to the first point and the second point.
9. One or more computer-readable non-transitory storage media, said one or more computer-readable non-transitory storage media comprising software, said software being operable to: Receive an image of a portion of the first user's face, wherein, The image was captured by a camera coupled to the first artificial reality head-mounted device worn by the first user; Access represents a three-dimensional 3D facial model of the first user's face; Identify one or more facial features depicted in the image; The relative pose of the camera of the first artificial reality head-mounted device with respect to the 3D facial model is determined by comparing the identified one or more facial features in the image with predetermined feature positions on the 3D facial model. The mapping relationship between the image of the first user's face and the 3D facial model is determined by projecting an image of a portion of the first user's face from the relative pose of the camera of the first artificial reality head-mounted device onto the 3D facial model. A reprojection package is generated using a reprojection module. The reprojection package includes the mapping relationship, the 3D facial model of the first user, and instructions for rendering the facial representation of the first user based on the mapping relationship and the 3D facial model. The first artificial reality head-mounted device sends the reprojection packet to the second artificial reality head-mounted device; as well as An output image of the facial representation of the first user is rendered from the viewpoint of the second user using the second artificial reality head-mounted device relative to the first user, wherein the output image is rendered using at least the 3D facial model and the mapping relationship between the image and the 3D facial model.
10. The medium according to claim 9, wherein, When the software is executed, it can also operate as follows: Receive multiple images corresponding to a second portion of the face of the first user, wherein the multiple images are captured by multiple cameras coupled to the first artificial reality head-mounted device worn by the first user; Synthesis is performed using a machine learning model to generate a synthetic image corresponding to the second portion of the first user's face; and A second mapping relationship between the synthetic image and the 3D facial model is determined by projecting the synthetic image of the second portion of the first user's face from a predetermined camera pose onto the 3D facial model.
11. The medium according to claim 9 or 10, wherein, The 3D facial model is a pre-defined 3D facial model representing multiple faces of multiple users.
12. The medium according to claim 11, wherein, When the software is executed, it can also operate as follows: The predetermined 3D facial model representing the multiple faces of the multiple users is deformed based at least on the one or more facial features identified in the image.
13. The medium according to claim 11, wherein, Determining the relative pose of the camera includes comparing the position of one or more facial features identified in the image with the position of the predetermined feature.
14. The medium according to claim 9 or 10, wherein, When the software is executed, it can also operate as follows: Based at least on the identified one or more facial features in the image, a predetermined 3D facial model representing multiple faces of multiple users is deformed.
15. The medium according to claim 9 or 10, wherein, Determining the relative pose of the camera includes comparing the position of one or more facial features identified in the image with the position of the predetermined feature.
16. A system comprising: One or more processors; as well as One or more computer-readable non-transitory storage media are coupled to one or more of the one or more processors and include instructions that, when executed by one or more of the one or more processors, are operable to cause the system to: Receive an image of a portion of a first user's face, wherein the image is captured by a camera coupled to a first artificial reality head-mounted device worn by the first user; Access represents a three-dimensional 3D facial model of the first user's face; Identify one or more facial features depicted in the image; The relative pose of the camera of the first artificial reality head-mounted device with respect to the 3D facial model is determined by comparing the identified one or more facial features in the image with predetermined feature positions on the 3D facial model. The mapping relationship between the image of the portion of the first user's face and the 3D facial model is determined by projecting the image of the portion of the first user's face from the relative pose of the camera of the first artificial reality head-mounted device onto the 3D facial model. A reprojection package is generated using a reprojection module. The reprojection package includes the mapping relationship, the 3D facial model of the first user, and instructions for rendering the facial representation of the first user based on the mapping relationship and the 3D facial model. The first AI reality headset sends the reprojection packet to the second AI reality headset; and An output image of the facial representation of the first user is rendered from the viewpoint of the second user using the second artificial reality head-mounted device relative to the first user, wherein the output image is rendered using at least the 3D facial model and the mapping relationship between the image and the 3D facial model.
17. The system according to claim 16, wherein, The one or more processors, when executing the instructions, are also capable of operating to: Receive multiple images corresponding to a second portion of the face of the first user, wherein the multiple images are captured by multiple cameras coupled to the first artificial reality head-mounted device worn by the first user; Synthesis is performed using a machine learning model to generate a synthetic image corresponding to the second portion of the first user's face; and A second mapping relationship between the synthetic image and the 3D facial model is determined by projecting the synthetic image of the second portion of the first user's face from a predetermined camera pose onto the 3D facial model.
18. The system according to claim 16 or 17, wherein, The 3D facial model is a pre-defined 3D facial model representing multiple faces of multiple users.
19. The system according to claim 18, wherein, The one or more processors, when executing the instructions, are also capable of operating to: The predetermined 3D facial model representing the multiple faces of the multiple users is deformed based at least on the one or more facial features identified in the image.
20. The system according to claim 18, wherein, Determining the relative pose of the camera includes comparing the position of one or more facial features identified in the image with the position of the predetermined feature.
21. The system according to claim 16 or 17, wherein, The one or more processors, when executing the instructions, are also capable of operating to: Based at least on the identified one or more facial features in the image, a predetermined 3D facial model representing multiple faces of multiple users is deformed.
22. The system according to claim 16 or 17, wherein, Determining the relative pose of the camera includes comparing the position of one or more facial features identified in the image with the position of the predetermined feature.
Citation Information
Patent Citations
Method and system of providing user facial displays in virtual or augmented reality for face occluding head mounted displays
US20180158246A1
Systems and methods for determining the scale of human anatomy from images
US20180336737A1
Detecting respiratory tract infection based on changes in coughing sounds
US20200245873A1