Image capture in extended reality environments
XR systems address the challenge of capturing high-quality self-images by generating and overlaying digital representations of users' poses onto real-world frames, enhancing the accuracy and realism of self-portraits without manual intervention.
Patent Information
- Application Number
- JP2024071505
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-02-11
- Filing Date
- 2024-04-25
- Publication Date
- 2025-12-03
- Estimated Expiration
- 2042-01-21
AI Technical Summary
Existing extended reality (XR) systems face challenges in capturing high-quality self-images or videos due to inward-facing cameras' limited field of view and the need for manual positioning of avatars, resulting in unnatural or inaccurate representations.
The XR system captures a user's pose and generates a digital representation, such as an avatar, which is then overlaid onto frames of the real-world environment, reflecting the user's pose and location, allowing for automatic and accurate self-image capture.
Enables high-fidelity, automatic, and realistic self-image capture within XR environments, reducing the need for manual user input and improving the quality of self-portraits.
Smart Images

Figure 0007779946000003 
Figure 0007779946000004 
Figure 0007779946000005
Abstract
Description
[Technical Field]
[0001] The present disclosure relates generally to techniques and systems for capturing images (such as self-images or "selfies") within extended reality environments. [Background technology]
[0002] Extended reality technology can be used to present virtual content to a user and / or combine a real environment from the physical world with a virtual environment to provide the user with an extended reality experience. The term extended reality can encompass virtual reality, augmented reality, mixed reality, etc. Extended reality systems can overlay virtual content onto an image of a real-world environment, allowing a user to experience an extended reality environment, which can be viewed by the user through an extended reality device (e.g., a head-mounted display, extended reality glasses, or other device). The real-world environment can include physical objects, people, or other real-world objects. XR technology can be implemented in a variety of applications and fields, including entertainment (e.g., games), teleconferencing, and education, among other applications and fields. Today, XR systems are being developed to provide users with the ability to capture photos or videos of themselves (e.g., "selfies").
[0003] Some types of devices (such as mobile phones and tablets) are equipped with mechanisms for users to capture images of themselves. However, self-image capture can be a challenging task for XR systems. For example, head-worn XR devices (e.g., head-mounted displays, XR glasses, and other devices) generally include cameras configured to capture scenes of a real-world environment. These cameras are positioned in an outward-facing direction (e.g., facing away from the user), and therefore, the cameras cannot capture images of the user. Furthermore, cameras positioned inside the XR device (e.g., inward-facing cameras) may not be able to capture more than a small portion of the user's face or head. Several XR systems have been developed to facilitate self-image capture. These XR systems can overlay a user's avatar on images and / or videos of the real-world environment. However, these avatars may have limited poses and / or expressions. Furthermore, the user may be required to manually position the avatar in the images and / or videos. Thus, there is a need for an improved XR system for self-image capture. Summary of the Invention [Means for solving the problem]
[0004] Systems and techniques that can be implemented to capture self-images within an extended reality environment are described herein. According to at least one example, an apparatus for capturing self-images within an extended reality environment is provided. The exemplary apparatus can include a memory (or multiple memories) and one or more processors (e.g., implemented in circuitry) coupled to the memory (or multiple memories). The processor (or multiple processors) is configured to: capture a pose of a user of the extended reality system, the user's pose including a location of the user within a real-world environment associated with the extended reality system; generate a digital representation of the user, the digital representation of the user reflecting the user's pose; capture one or more frames of the real-world environment; and overlay the digital representation of the user on the one or more frames of the real-world environment.
[0005] Another example apparatus may include means for capturing a pose of a user of an extended reality system, the user's pose including a location of the user in a real-world environment associated with the extended reality system; means for generating a digital representation of the user, the digital representation of the user reflecting the user's pose; means for capturing one or more frames of the real-world environment; and means for overlaying the digital representation of the user onto the one or more frames of the real-world environment.
[0006] In another example, a method for capturing a self-portrait within an extended reality environment is provided. The exemplary method can include capturing a pose of a user of an extended reality system, the user's pose including the user's location within a real-world environment associated with the extended reality system. The method can also include generating a digital representation of the user, the digital representation of the user reflecting the user's pose. The method can include capturing one or more frames of the real-world environment. The method can further include overlaying the digital representation of the user on the one or more frames of the real-world environment.
[0007] In some aspects, the method may be performed by an extended reality system. In such aspects, the method may include capturing a pose of a user of the extended reality system by the extended reality system, the user's pose including the user's location in a real-world environment associated with the extended reality system, generating a digital representation of the user by the extended reality system, the digital representation of the user reflecting the user's pose, capturing one or more frames of the real-world environment by the extended reality system, and overlaying the digital representation of the user onto the one or more frames of the real-world environment by the extended reality system for presentation to the user using the extended reality system.
[0008] In another example, a non-transitory computer-readable medium for capturing self-portraits within an extended reality environment is provided. The exemplary non-transitory computer-readable medium can store instructions that, when executed by one or more processors, cause the one or more processors to: capture a pose of a user of an extended reality system, the user's pose including a location of the user within a real-world environment associated with the extended reality system; generate a digital representation of the user, the digital representation of the user reflecting the user's pose; capture one or more frames of the real-world environment; and overlay the digital representation of the user on the one or more frames of the real-world environment.
[0009] In some aspects, overlaying the digital representation of the user onto one or more frames of the real-world environment may include overlaying the digital representation of the user within a frame location corresponding to the user's location within the real-world environment.
[0010] In some aspects, generating the digital representation of the user may be performed before capturing one or more frames of the real-world environment. In some examples, the methods, devices, and computer-readable media described above can include displaying the digital representation of the user in a display location corresponding to the user's location in the real-world environment within a display of the extended reality system through which the real-world environment is viewed. In one example, the methods, devices, and computer-readable media described above can include detecting user input corresponding to instructions to capture one or more frames of the real-world environment while the digital representation of the user is displayed within a display of the extended reality system, and capturing one or more frames of the real-world environment based on the user input.
[0011] In some aspects, capturing one or more frames of the real-world environment may be performed before capturing the user's pose. In some examples, the methods, devices, and computer-readable media described above may include displaying a digital representation of the user in a display location corresponding to the user's location in the real-world environment within a display of the extended reality system on which the one or more frames of the real-world environment are displayed. In one example, the methods, devices, and computer-readable media described above may include updating the display location of the digital representation of the user based on detecting a change in the user's location within the real-world environment. In another example, the methods, devices, and computer-readable media described above may include detecting a user input corresponding to an instruction to capture a pose of the user while the digital representation of the user is displayed within a display of the extended reality system, and capturing the user's pose based on the user input.
[0012] In some aspects, the first digital representation is based on a first machine learning algorithm and the second digital representation of the user is based on a second machine learning algorithm. For example, generating the digital representation of the user can include generating the first digital representation of the user based on a first machine learning algorithm and obtaining the second digital representation of the user based on a second machine learning algorithm. In some aspects, the second digital representation of the user can be a higher-fidelity digital representation of the user than the first digital representation of the user. For example, in some examples, the methods, apparatus, and computer-readable media described above can include generating a first digital representation of the user of a first fidelity and obtaining a second digital representation of the user of a second fidelity, the second fidelity being higher than the first fidelity. In some examples, the methods, devices, and computer-readable media described above may include displaying a first digital representation of a user within a display of an extended reality system before a pose of the user is captured (e.g., using the user's current pose), generating a second digital representation of the user based on the captured pose of the user, and overlaying the second digital representation of the user onto one or more frames of a real-world environment. In other examples, the methods, devices, and computer-readable media described above may include displaying a first digital representation of a user within a display of an extended reality system before one or more frames of a real-world environment are captured (e.g., using the user's captured pose), generating a second digital representation of the user based on the captured one or more frames of the real-world environment, and overlaying the second digital representation of the user onto one or more frames of the real-world environment. In some aspects, generating the first digital representation of the user may include implementing a first machine learning algorithm on the extended reality system.In some examples, obtaining the second digital representation of the user may include causing a server configured to generate digital representations of users to generate the second digital representation of the user based on implementing a second machine learning algorithm.
[0013] In some aspects, the methods, apparatus, and computer-readable media described above may include capturing a pose of a person in a real-world environment, generating a digital representation of the person, where the digital representation of the person reflects the pose of the person, and overlaying the digital representation of the user and the digital representation of the person onto one or more frames of the real-world environment. In some examples, the digital representation of the person may be generated based at least in part on information associated with the digital representation of the person received from the person's extended reality system. In some aspects, the information associated with the digital representation of the person may include a machine learning model trained to generate the digital representation of the person.
[0014] In some aspects, the methods, apparatus, and computer-readable media described above may include capturing a plurality of poses of a user associated with a plurality of frames, generating a plurality of digital representations of the user corresponding to the plurality of frames, and overlaying the plurality of digital representations of the user onto one or more frames of a real-world environment, wherein the one or more frames of the real-world environment include the plurality of frames of the real-world environment.
[0015] In some aspects, the methods, apparatus, and computer-readable media described above may include generating a digital representation of a user using a first machine learning algorithm and overlaying the digital representation of the user onto one or more frames of a real-world environment using a second machine learning algorithm.
[0016] In some aspects, capturing the user's pose can include capturing image data using an inward-facing camera system of the extended reality system. In some examples, capturing the user's pose can include determining a facial expression of the user. In other examples, capturing the user's pose can include determining a gesture of the user.
[0017] In some aspects, the methods, apparatus, and computer-readable media described above may include determining a location of a user within a real-world environment based at least in part on generating a three-dimensional map of the real-world environment.
[0018] In some aspects, capturing one or more frames of the real-world environment may include capturing the image data using an outward-facing camera system of the extended reality system.
[0019] In some aspects, each of the devices described above is or includes a camera, a mobile device (e.g., a mobile phone or so-called "smartphone," or other mobile device), a smart wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a server computer, a vehicle (e.g., an autonomous vehicle), or other device. In some aspects, the device includes one or more cameras for capturing one or more videos and / or images. In some aspects, the device further includes a display for displaying the one or more videos and / or images. In some aspects, the devices described above may include one or more sensors.
[0020] This Summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used alone to determine the scope of the claimed subject matter, which subject matter should be understood by reference to the entire specification of this patent, any or all drawings, and appropriate portions of each claim.
[0021] The above, together with other features and examples, will become more apparent with reference to the following specification, claims, and accompanying drawings.
[0022] Illustrative examples of the present application are described in detail below with reference to the following figures: [Brief explanation of the drawings]
[0023] [Figure 1A] 1 is a diagram of an exemplary avatar generated by an extended reality system, according to some examples. [Figure 1B] 1 is a diagram of an exemplary avatar generated by an extended reality system, according to some examples. [Figure 1C] 1 is a diagram of an exemplary avatar generated by an extended reality system, according to some examples. [Figure 1D] 1 is a diagram of an exemplary avatar generated by an extended reality system, according to some examples. [Figure 2] FIG. 1 is a block diagram illustrating an example architecture of an extended reality system, according to some examples. [Figure 3A] FIG. 1 is a block diagram of an exemplary system for self-image capture within an extended reality environment, according to some examples. [Figure 3B] FIG. 1 is a block diagram of an exemplary system for self-image capture within an extended reality environment, according to some examples. [Figure 4A]1 is a diagram of an exemplary self-image generated by a system for self-image capture within an extended reality environment, according to some examples. [Figure 4B] 1 is a diagram of an exemplary self-image generated by a system for self-image capture within an extended reality environment, according to some examples. [Figure 4C] 1 is a diagram of an exemplary self-image generated by a system for self-image capture within an extended reality environment, according to some examples. [Figure 4D] 1 is a diagram of an exemplary self-image generated by a system for self-image capture within an extended reality environment, according to some examples. [Figure 4E] 1 is a diagram of an exemplary self-image generated by a system for self-image capture within an extended reality environment, according to some examples. [Figure 4F] 1 is a diagram of an exemplary self-image generated by a system for self-image capture within an extended reality environment, according to some examples. [Figure 5A] 1 is a flow diagram illustrating an example of a process for self-video capture within an extended reality environment. [Figure 5B] 1 is a flow diagram illustrating an example of a process for self-video capture within an extended reality environment. [Figure 6A] 1 is a flow diagram illustrating an example of a process for multi-user selfie image capture within an extended reality environment. [Figure 6B] 1 is a flow diagram illustrating an example of a process for multi-user selfie image capture within an extended reality environment. [Figure 6C] 1 is a block diagram illustrating an example system for multi-user self-image capture within an extended reality environment, according to some examples. [Figure 7]1 is a flow diagram of a process for self-image capture within an extended reality environment, according to some examples. [Figure 8] FIG. 1 illustrates an example of a deep learning neural network, according to some examples. [Figure 9] FIG. 1 illustrates an example of a convolutional neural network, according to some examples. [Figure 10] FIG. 1 illustrates an example system for implementing certain aspects described herein. DETAILED DESCRIPTION OF THE INVENTION
[0024] Some aspects and examples of the present disclosure are provided below. As will be apparent to those skilled in the art, some of these aspects and examples may be applied independently, and some of them may be applied in combination. In the following description, for purposes of explanation, specific details are set forth in order to provide a thorough understanding of the subject matter of the present application. However, it will be apparent that various examples may be practiced without these specific details. The figures and descriptions are not intended to be limiting.
[0025] The following description provides illustrative examples only and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the following description will provide those skilled in the art with an enabling description for implementing the illustrative examples. It should be understood that various changes may be made in the function and arrangement of elements without departing from the spirit and scope of the present application as set forth in the appended claims.
[0026] As described above, extended reality (XR) technology can be used to present virtual content to a user and / or combine a real environment from the physical world with a virtual environment to provide the user with an extended reality experience. The term extended reality can encompass virtual reality (VR), augmented reality (AR), mixed reality (MR), etc. XR systems can overlay virtual content onto an image of a real-world environment, allowing a user to experience an extended reality environment, which can be viewed by the user through an extended reality device (e.g., a head-mounted display, extended reality glasses, or other device). The real-world environment can include physical objects, people, or other real-world objects. XR technology can be implemented in a variety of applications and fields, including entertainment (e.g., games), room design and / or virtual shopping, food / health monitoring, video calling, teleconferencing, and education, among other applications and fields.
[0027] Some types of devices (such as mobile phones and tablets) are equipped with mechanisms for users to capture images of themselves (e.g., "selfie images" or "selfies"). However, selfie image capture can be a challenging task for XR systems. For example, XR devices worn on a user's head (e.g., head-mounted displays, XR glasses, and other devices) generally include cameras configured to capture scenes of the real-world environment. These cameras are positioned in an outward-facing direction (e.g., pointed away from the user), and therefore the cameras cannot capture images of the user. Furthermore, cameras positioned inside the XR device (e.g., inward-facing cameras) may not be able to capture more than a small portion of the user's face or head. Several XR systems have been developed to facilitate selfie image capture. These XR systems can overlay a user's avatar on images and / or videos of the real-world environment. However, the avatar may have limited poses and / or facial expressions. Additionally, the user may be required to manually position the avatar within the image and / or video, which may result in an unnatural (e.g., unrealistic) image and / or inaccurate framing of the avatar relative to the background.
[0028] This disclosure describes systems, apparatuses, methods, and computer-readable media (collectively referred to as “systems and techniques”) for performing image capture within an XR environment. In some aspects, the techniques and systems provide the ability for a head-mounted XR system to capture one or more self-images (e.g., photographs and / or videos) of a user wearing the head-mounted XR system. For example, the XR system can overlay a digital representation of the user (e.g., an avatar or other type of digital representation) over one or more frames captured by the XR system's front-facing camera. The digital representation of the user can reflect the user's pose (e.g., the user's gestures, limb positions, facial expressions, etc.). The XR system can overlay a digital representation of the user within the frame such that the location of the digital representation corresponds to the user's actual location within the real-world environment (e.g., the user's location when the user's pose was captured). When the digital representation is overlaid within the frame, the frame may appear as if the user (e.g., the user's digital representation) was facing the XR system's camera when the frame was captured. In this way, the frame corresponds to the user's self-image (or “selfie”).
[0029] In some cases, the process of generating a user's self-image may involve at least two operations. One operation involves capturing one or more frames to be used as a background for the self-image. Another operation involves capturing a user's pose to be incorporated into a digital representation of the user (e.g., an avatar or other digital representation). In some cases, the XR system may capture the user's pose through one or more tracking capabilities of the XR system. The tracking capabilities of the XR system may include mechanisms and / or techniques for gaze tracking, six degrees of freedom (6DOF) position tracking, hand tracking, body tracking, facial expression tracking, any combination thereof, and / or other types of tracking. In some examples, capturing the user's pose and incorporating the user's pose into a digital representation (e.g., an avatar) may be referred to as "puppeteering" the user's digital representation. The two operations for generating the self-image may be performed sequentially (e.g., one operation after the other). Furthermore, the operations may be performed in any order. For example, a background frame may be captured before the user's pose is captured, or the user's pose may be captured before the background frame is captured. The specific process for implementing the actions can vary based on the order of the actions. In some cases, after detecting user input corresponding to initiating self-image capture mode, the XR system can prompt the user to perform the actions in a certain order. In other cases, the XR system can detect user input corresponding to a preferred order.
[0030] In a first use case scenario, a digital representation (e.g., an avatar) is manipulated before a background frame is captured. In this scenario, the XR system can capture a pose of the user and generate a digital representation of the user that reflects the user's captured pose. In some cases, the XR system can represent the digital representation within the XR system's display. For example, the XR system can position the digital representation within the display so that the digital representation appears to be located in the same location in the real-world environment as the user was when the user's state was captured. As the user moves within the real-world environment, portions of the real-world environment (e.g., a scene) seen through the XR system's display change, but the digital representation appears to remain in the same location. In this way, the user's digital representation is "world-locked." To determine the user's location when the user's state was captured, the XR system can map the environment using one or more mapping and / or localization techniques, including computer vision-based tracking, model-based tracking, simultaneous localization and mapping (SLAM) techniques, any combination thereof, and / or other mapping and / or localization techniques. After the XR system generates an avatar or other digital representation of the user (the second operation above), the XR system can capture a frame that will become the background of the self-portrait (the first operation above). For example, the user can move through an environment while the digital representation (e.g., an avatar) is presented on the XR system's display. When the user determines that the view through the display is an appropriate and / or desirable background for the self-portrait, the user can instruct the XR system to capture a frame of the environment. The XR system can then generate a composite frame in which the user's digital representation is overlaid on the captured frame of the environment.
[0031] In a second use case scenario, a background frame is captured before the digital representation of the user is manipulated. In this scenario, the XR system can capture frames of the real-world environment at the user's direction. In some cases, the background frame is represented within a portion of the XR system's display (e.g., within a preview window). After the background frame is captured and / or displayed within the preview window (the first action above), the XR system can manipulate the user's avatar or other digital representation based on the user's pose (the second action above). For example, the XR system can represent the avatar within the preview window (e.g., at a predetermined location). As the user moves through the environment, the XR system can move the avatar to a corresponding location within the preview window. In this way, the avatar is "headlocked." When the user is satisfied with their location and / or pose, the user can instruct the XR system to capture the user's pose. The XR system can then generate a composite frame in which the user's digital representation (e.g., avatar) is overlaid on the captured frame of the environment.
[0032] In a third use case scenario, the XR system can generate a self-image (e.g., a video) having multiple frames. In one example of this scenario, the XR system can generate a video by manipulating a digital representation of the user (e.g., an avatar representing the user) before capturing background frames. For example, the XR system can capture a pose of the user for a period of time and generate a digital representation of the user (e.g., an avatar) that reflects any changes in the user's pose during that period of time. The period of time can correspond to a predetermined period of time, a predetermined number of frames, or a period of time specified by the user. The XR system can then record frames of the real-world environment during the same period of time. The scene of the real-world environment can remain constant throughout the period of time, or the scene can change (e.g., the user can move within the environment between frames being captured). In some cases, the XR system can present a digital representation of the user within the XR system's display while the XR system is recording frames of the environment. For example, the XR system may represent a digital representation of a user so that the digital representation appears to be in the same location as the user when the user's pose was captured (e.g., the digital representation of the user is world-locked). In another example of this use case scenario, the XR system may generate a video by capturing background frames before manipulating the digital representation of the user. For example, the XR system may capture frames of an environment that will become a static background for the video. The XR system may then capture the user's pose for a period of time and generate a digital representation of the user (e.g., an avatar) that reflects any changes in the user's pose during that period of time. As the user moves within the environment during that period of time, the XR system may display the digital representation of the user along with the corresponding movement and / or pose in a preview window that contains the frames. In this example, the digital representation of the user (e.g., an avatar representing the user) is head-locked.
[0033] In a fourth use case scenario, an XR system can generate a selfie image that includes multiple people. In one example of this scenario, the XR system can capture a user's pose while the user poses with the one or more people. The XR system can then capture a background frame that includes the one or more people. For example, the user can move within the environment to face one or more people while the one or more people remain in their original pose. The XR system can then insert a digital representation (e.g., an avatar) that reflects the user's state into the background frame. In another example of this use case scenario, each person to be included in the selfie image is wearing a head-mounted XR system. In this example, each XR system can capture the pose of the corresponding person and generate a digital representation (e.g., an avatar) that reflects the person's pose. In addition, one of the XR systems can capture a background frame. The XR system that captures the background frame can send a request to the other XR systems for the digital representation (e.g., avatar) generated by the other XR system. The XR system can then combine the background frame and the digital representation. In this example, the background frame may be captured before the digital representation is generated, or the digital representation may be generated before the background frame is captured.
[0034] An XR system can generate an avatar or other digital representation of a user in various ways. In some examples, the XR system can capture one or more frames of a user (e.g., one or more full-body images of the user). The frames can be input into a model that generates a digital representation based on physical features detected in the frames. In some cases, the model can be trained to adapt to the user's digital representation based on the user's captured pose. For example, if the XR system determines that a user is waving, the model can generate a digital representation that resembles the user waving, even if the model does not have an actual image of the user waving. In one example, the model can be a machine learning model such as a neural network. In a non-limiting example, the model can be a generative adversarial network (GAN). In some cases, the XR system can generate an initial digital representation (e.g., an avatar) of the user. For example, the initial digital representation of the user can be presented within the XR system's display while the XR system captures background frames and / or the user's pose. Once the user's pose and / or placement of the digital representation within the background frame is confirmed, the XR system may replace the initial digital representation with a digital representation of the user (e.g., a final avatar for the user). In some cases, the final digital representation may be of higher quality (e.g., more realistic) than the initial digital representation. For example, the initial avatar may appear cartoon-like, while the final avatar may be photorealistic (or nearly photorealistic). Because generating a photorealistic digital representation may involve a large amount of time and / or processing power, using a relatively lower-quality initial digital representation may reduce the latency and / or workload of the XR system during the process of generating the selfie-image.
[0035] Various aspects of the application are described with reference to the figures. FIGS. 1A, 1B, and 1C provide illustrations of example avatars that may be generated by some existing XR systems. For example, FIG. 1A shows a selfie-picture frame 102 including an avatar 108 overlaid on a background frame 106. In one example, the XR system can generate the avatar 108 by inputting one or more images of the user into a model (e.g., a machine learning-based model) trained to generate an avatar whose physical appearance corresponds to the user's physical appearance. As shown, the avatar 108 may be a cartoon-like or other abstract representation of the user (rather than a photorealistic or near-photorealistic representation of the user). In some cases, the XR system may request and / or prompt the user to position the avatar 108 at a selected location within the background frame 106. For example, the XR system may provide a positioning tool 110 shown in FIG. 1A that allows the user to reposition the avatar 108 by dragging the positioning tool 110 within the background frame 106. The selfie image frame 104 shown in FIG. 1B shows the avatar 108 after it has been positioned using a positioning tool 110. In some cases, the selfie image capture system may allow a user to customize the selfie image frame 104 by selecting a pose (e.g., a gesture, an emotion, a facial expression, a movement, etc.) for the avatar 108. For example, the selfie image capture system may provide a menu or list of predetermined and / or pre-configured poses and detect user input corresponding to a selection. Further, in some cases, the selfie image capture system may allow a user to animate the avatar 108 within the background frame 106. For example, as shown in FIG. 1C , the selfie image capture system may allow a user to draw a path for the avatar 108 using the positioning tool 110. Based on the user input corresponding to the path, the selfie image capture system may generate a video showing the avatar 108 moving along the path.1D is an illustration of another exemplary avatar that may be generated by some existing XR systems. For example, FIG. 1D shows frame 116 including two avatars (e.g., avatars 118 and 120). In this example, users corresponding to avatars 118 and 120 can view both avatars within the display of an XR device (e.g., an HMD) worn by the user.
[0036] In some cases, XR systems configured to generate the avatars shown in FIGS. 1A-1D may have various limitations that make the XR system unsuitable and / or undesirable for generating self-portraits of users wearing HMDs or other types of XR devices. For example, the XR system may involve (e.g., require) manual user input to position the avatar within the frame. As a result, the frame including the rendered avatar may not appear like a natural or realistic "selfie." Furthermore, the rendered avatar may not accurately reflect the current (e.g., actual) user's pose. Moreover, due to the large computing resources involved in rendering the avatar, avatars generated by these XR systems may be of low quality (e.g., low fidelity).
[0037] FIG. 2 illustrates an exemplary extended reality system 200 according to some aspects of the present disclosure. The extended reality system 200 can be part of or implemented by a single computing device or multiple computing devices. In some cases, the extended reality system 200 can be part of or implemented by an XR system or device. For example, the extended reality system 200 can run (or execute) XR applications and implement XR actions. An XR system or device that includes and / or implements the extended reality system 200 can be an XR head-mounted display (HMD) device (e.g., a virtual reality (VR) headset, an augmented reality (AR) headset, or a mixed reality (MR) headset), XR glasses (e.g., AR glasses), among other XR systems or devices. In some examples, the extended reality system 200 may be part of or may be implemented by any other device or system, such as a camera system (e.g., digital camera, IP camera, video camera, security camera, etc.), a telephone system (e.g., smartphone, cellular phone, conferencing system, etc.), a desktop computer, a laptop or notebook computer, a tablet computer, a set-top box, a network-connected television (or so-called "smart" television), a display device, a gaming console, a video streaming device, an Internet of Things (IoT) device, a vehicle (or a computing device in a vehicle), and / or any other suitable electronic device.
[0038] In some examples, the extended reality system 200 can perform tracking and localization, mapping of the physical world (e.g., a scene), and positioning and rendering of virtual content on a display (e.g., a screen, a viewable plane / area, and / or another display) as part of an XR experience. For example, the extended reality system 200 can generate a map (e.g., a 3D map) of the scene in the physical world, track the pose (e.g., location and position) of the extended reality system 200 relative to the scene (e.g., relative to the 3D map of the scene), position and / or anchor virtual content at specific locations on the map of the scene, and render the virtual content on the display. The extended reality system 200 can render virtual content on the display such that the virtual content appears to be at a location in the scene that corresponds to the specific location on the map of the scene where the virtual content is positioned and / or anchored. In some examples, the display can include glasses, a screen, lenses, and / or other display mechanisms that allow a user to view the real-world environment and also allow XR content to be displayed thereon.
[0039] 2, extended reality system 200 may include one or more image sensors 202, accelerometers 204, gyroscopes 206, storage 208, computational components 210, an XR engine 220, a self-image engine 222, an image processing engine 224, and a rendering engine 226. It should be noted that the components 202-226 shown in FIG. 2 are non-limiting examples provided for illustrative and descriptive purposes, and that other examples may include more, fewer, or different components than those shown in FIG. 2. For example, in some cases, extended reality system 200 may include one or more other sensors (e.g., one or more inertial measurement units (IMUs), radar, light detection and ranging (LIDAR) sensors, audio sensors, etc.), one or more display devices, one or more other processing engines, one or more other hardware components, and / or one or more other software and / or hardware components not shown in FIG. 2. An example architecture and example hardware components that may be implemented by the extended reality system 200 are further described below with respect to FIG.
[0040] For simplicity and purposes of explanation, one or more image sensors 202 are referred to herein as image sensor 202 (e.g., in the singular). However, those skilled in the art will recognize that extended reality system 200 can include a single image sensor or multiple image sensors. Additionally, reference to any of the components (e.g., 202-226) of extended reality system 200 in the singular or plural should not be construed as limiting the number of such components implemented by extended reality system 200 to one or more. For example, reference to accelerometer 204 in the singular should not be construed as limiting the number of accelerometers implemented by extended reality system 200 to one. Those skilled in the art will recognize that for any of components 202-226 shown in FIG. 2, extended reality system 200 can include only one of such components or multiple of such components.
[0041] The extended reality system 200 may include or communicate (wired or wirelessly) with input devices. The input devices may include any suitable input devices, such as a touchscreen, a pen or other pointer device, a keyboard, a mouse, buttons or keys, a microphone for receiving voice commands, a gesture input device for receiving gesture commands, any combination thereof, and / or other input devices. In some cases, the image sensor 202 may capture images that can be processed to interpret the gesture commands.
[0042] The extended reality system 200 can be part of or implemented by a single computing device or multiple computing devices. In some examples, the extended reality system 200 can be part of an electronic device (or multiple devices), such as an extended reality head-mounted display (HMD) device, extended reality glasses (e.g., augmented reality or AR glasses), a camera system (e.g., a digital camera, an IP camera, a video camera, a security camera, etc.), a telephone system (e.g., a smartphone, a cellular phone, a conferencing system, etc.), a desktop computer, a laptop or notebook computer, a tablet computer, a set-top box, a smart television, a display device, a game console, a video streaming device, an Internet of Things (IoT) device, and / or any other suitable electronic device.
[0043] In some implementations, one or more of the image sensor 202, the accelerometer 204, the gyroscope 206, the storage 208, the computational component 210, the XR engine 220, the self-image engine 222, the image processing engine 224, and the rendering engine 226 can be part of the same computing device. For example, in some cases, the one or more of the image sensor 202, the accelerometer 204, the gyroscope 206, the storage 208, the computational component 210, the XR engine 220, the self-image engine 222, the image processing engine 224, and the rendering engine 226 can be integrated into an HMD, extended reality glasses, a smartphone, a laptop, a tablet computer, a gaming system, and / or any other computing device. However, in some implementations, one or more of the image sensor 202, accelerometer 204, gyroscope 206, storage 208, computational component 210, XR engine 220, self-image engine 222, image processing engine 224, and rendering engine 226 can be part of two or more separate computing devices. For example, in some cases, some of the components 202-226 can be part of or implemented by one computing device, and the remaining components can be part of or implemented by one or more other computing devices.
[0044] Storage 208 can be any storage device for storing data. Moreover, storage 208 can store data from any of the components of extended reality system 200. For example, storage 208 can store data from image sensor 202 (e.g., image or video data), data from accelerometer 204 (e.g., measurements), data from gyroscope 206 (e.g., measurements), data from computation component 210 (e.g., processing parameters, preferences, virtual content, rendering content, scene maps, tracking and localization data, object detection data, privacy data, XR application data, facial recognition data, occlusion data, etc.), data from XR engine 220, data from self-image engine 222, data from image processing engine 224, and / or data (e.g., output frames) from rendering engine 226. In some examples, storage 208 can include a buffer for storing frames for processing by computation component 210.
[0045] The one or more computing components 210 may include a central processing unit (CPU) 212, a graphics processing unit (GPU) 214, a digital signal processor (DSP) 216, and / or an image signal processor (ISP) 218. The computing component 210 may perform various operations such as image enhancement, computer vision, graphics rendering, extended reality (e.g., tracking, localization, pose estimation, mapping, content fixation, content rendering, etc.), image / video processing, sensor processing, recognition (e.g., text recognition, face recognition, object recognition, feature recognition, tracking or pattern recognition, scene recognition, occlusion detection, etc.), machine learning, filtering, and any of the various operations described herein. In this example, the computing component 210 implements an XR engine 220, a self-image engine 222, an image processing engine 224, and a rendering engine 226. In other examples, the computing component 210 may also implement one or more other processing engines.
[0046] Image sensor 202 can include any image and / or video sensor or capture device. In some examples, image sensor 202 can be part of a multiple camera assembly, such as a dual camera assembly. Image sensor 202 can capture images and / or video content (e.g., raw images and / or video data), which can then be processed by computation component 210, XR engine 220, self-image engine 222, image processing engine 224, and / or rendering engine 226, as described herein.
[0047] In some examples, the image sensor 202 can capture image data, generate frames based on the image data, and / or provide the image data or frames to the XR engine 220, the self-image engine 222, the image processing engine 224, and / or the rendering engine 226 for processing. A frame can include a video frame of a video sequence or a still image. A frame can include a pixel array representing a scene. For example, a frame can be a red-green-blue (RGB) frame having red, green, and blue color components per pixel, a luma-red-difference-blue-difference (YCbCr) frame having a luma component and two chroma (color) components (red difference and blue difference) per pixel, or any other suitable type of color or monochrome picture.
[0048] In some cases, image sensor 202 (and / or other image sensors or cameras of extended reality system 200) can be configured to also capture depth information. For example, in some implementations, image sensor 202 (and / or other cameras) can include an RGB-depth (RGB-D) camera. In some cases, extended reality system 200 can include one or more depth sensors (not shown) that are separate from image sensor 202 (and / or other cameras) and can capture depth information. For example, such depth sensors can acquire depth information independently from image sensor 202. In some examples, the depth sensor can be physically located in the same general location as image sensor 202 but can operate at a different frequency or frame rate than image sensor 202. In some examples, the depth sensor can take the form of a light source that can project a structured or textured light pattern, which may include one or more narrow bands of light, onto one or more objects in a scene. Depth information can then be obtained by exploiting geometric distortions of the projected pattern caused by the surface shape of the object. In one example, depth information may be obtained from a stereo sensor, such as a combination of an infrared structured light projector and an infrared camera registered with a camera (e.g., an RGB camera).
[0049] As mentioned above, in some cases, extended reality system 200 may also include one or more sensors (not shown) other than image sensor 202. For example, the one or more sensors may include one or more accelerometers (e.g., accelerometer 204), one or more gyroscopes (e.g., gyroscope 206), and / or other sensors. The one or more sensors may provide velocity, orientation, and / or other position-related information to computation component 210. For example, accelerometer 204 may detect acceleration by extended reality system 200 and generate acceleration measurements based on the detected acceleration. In some cases, accelerometer 204 may provide one or more translation vectors (e.g., up / down, left / right, forward / backward) that may be used to determine the position or pose of extended reality system 200. Gyroscope 206 may detect and measure the orientation and angular velocity of extended reality system 200. For example, gyroscope 206 may be used to measure the pitch, roll, and yaw of extended reality system 200. In some cases, the gyroscope 206 can provide one or more rotation vectors (e.g., pitch, yaw, roll). In some examples, the image sensor 202 and / or the XR engine 220 can use measurements obtained by the accelerometer 204 (e.g., one or more translation vectors) and / or the gyroscope 206 (e.g., one or more rotation vectors) to calculate a pose of the extended reality system 200. As mentioned above, in other examples, the extended reality system 200 can also include other sensors, such as an inertial measurement unit (IMU), a magnetometer, a machine vision sensor, a smart scene sensor, a voice recognition sensor, an impact sensor, a shock sensor, a position sensor, a tilt sensor, etc.
[0050] In some cases, the one or more sensors may include at least one IMU, which is an electronic device that uses a combination of one or more accelerometers, one or more gyroscopes, and / or one or more magnetometers to measure specific forces, angular velocities, and / or orientations of extended reality system 200. In some examples, the one or more sensors may output measured information associated with the capture of images captured by image sensor 202 (and / or other cameras of extended reality system 200) and / or depth information obtained using one or more depth sensors of extended reality system 200.
[0051] The output of one or more sensors (e.g., accelerometer 204, gyroscope 206, one or more IMUs, and / or other sensors) can be used by extended reality engine 220 to determine the pose of extended reality system 200 (also referred to as head pose) and / or the pose of image sensor 202 (or other camera of extended reality system 200). In some cases, the pose of extended reality system 200 and the pose of image sensor 202 (or other camera) can be the same. The pose of image sensor 202 refers to the position and orientation of image sensor 202 relative to a frame of reference. In some implementations, the pose of a camera can be determined in six degrees of freedom (6DOF), which refer to three translational components (e.g., which may be given by X (horizontal), Y (vertical), and Z (depth) coordinates relative to a frame of reference such as an image plane) and three angular components (e.g., roll, pitch, and yaw relative to the same frame of reference).
[0052] In some cases, a device tracker (not shown) can use measurements from one or more sensors and image data from image sensor 202 to track the pose (e.g., a 6DOF pose) of extended reality system 200. For example, the device tracker can fuse visual data from the captured image data with inertial measurement data (e.g., using a visual tracking solution) to determine the position and motion of extended reality system 200 relative to the physical world (e.g., a scene) and a map of the physical world. As described below, in some examples, when tracking the pose of extended reality system 200, the device tracker can generate a three-dimensional (3D) map of the scene (e.g., the real world) and / or generate updates for the 3D map of the scene. The 3D map updates can include, for example, without limitation, new or updated features and / or feature or landmark points associated with the scene and / or the 3D map of the scene, localization updates that identify or update the position of extended reality system 200 within the scene and the 3D map of the scene, etc. A 3D map can provide a digital representation of a scene in the real / physical world. In some examples, the 3D map can anchor location-based objects and / or content to real-world coordinates and / or objects. The extended reality system 200 can use the mapped scene (e.g., a scene in the physical world represented by and / or associated with the 3D map) to merge the physical world with the virtual world and / or merge virtual content or objects with the physical environment.
[0053] In some aspects, the pose of the image sensor 202 and / or the extended reality system 200 as a whole can be determined and / or tracked by the computational component 210 using a visual tracking solution based on images captured by the image sensor 202 (and / or other cameras of the extended reality system 200). For example, in some examples, the computational component 210 can perform tracking using computer vision-based tracking, model-based tracking, and / or simultaneous localization and map building (SLAM) techniques. For example, the computational component 210 can perform SLAM or can communicate (wired or wirelessly) with a SLAM engine (not shown). SLAM refers to a class of techniques in which a map of an environment (e.g., a map of the environment modeled by the extended reality system 200) is created while simultaneously tracking the pose of the camera (e.g., the image sensor 202) and / or the extended reality system 200 relative to the map. The map can be referred to as a SLAM map and can be 3D. SLAM techniques can be implemented using color or grayscale image data captured by image sensor 202 (and / or other cameras of extended reality system 200) and can be used to generate estimates of 6DOF pose measurements of image sensor 202 and / or extended reality system 200. Such SLAM techniques configured to implement 6DOF tracking can be referred to as 6DOF SLAM. In some cases, the output of one or more sensors (e.g., accelerometer 204, gyroscope 206, one or more IMUs, and / or other sensors) can be used to estimate, correct, and / or otherwise adjust the estimated pose.
[0054] In some cases, 6DOF SLAM (e.g., 6DOF tracking) can associate observed features from several input images from the image sensor 202 (and / or other cameras) into a SLAM map. For example, 6DOF SLAM can use feature point associations from input images to determine the pose (position and orientation) of the image sensor 202 and / or the extended reality system 200 relative to the input images. 6DOF mapping can also be performed to update the SLAM map. In some cases, a SLAM map maintained using 6DOF SLAM can include 3D feature points triangulated from two or more images. For example, keyframes can be selected from the input images or video stream to represent the observed scene. For every keyframe, a respective 6DOF camera pose associated with the image can be determined. The pose of the image sensor 202 and / or the extended reality system 200 can be determined by projecting features from the 3D SLAM map onto the image or video frame and updating the camera pose from the verified 2D-3D correspondence.
[0055] In one illustrative example, the computation component 210 can extract feature points from all input images or from each keyframe. As used herein, a feature point (also referred to as a registration point) is a distinctive or identifiable part of an image, such as a part of a hand or the edge of a table, among others. The features extracted from a captured image can represent distinctive feature points along three-dimensional space (e.g., coordinates on the X, Y, and Z axes), and every feature point can have an associated feature location. The feature points in a keyframe either match (are the same as or correspond to) or do not match feature points in a previously captured input image or keyframe. Feature detection can be used to detect feature points. Feature detection can include image processing operations used to examine one or more pixels of an image to determine whether a feature is present at a particular pixel. Feature detection can be used to process the entire captured image or portions of an image. For each image or keyframe, once a feature is detected, a local image patch near the feature can be extracted. Features may be extracted using any suitable technique, such as Scale Invariant Feature Transform (SIFT) (which locates features and generates their descriptions), Speed Up Robust Features (SURF), Gradient Position Orientation Histogram (GLOH), Normalized Cross Correlation (NCC), or other suitable technique.
[0056] In some cases, the extended reality system 200 may also track the user's hands and / or fingers to allow the user to interact with and / or control virtual content in the virtual environment. For example, the extended reality system 200 may track the pose and / or movement of the user's hands and / or fingertips to identify or interpret user interactions with the virtual environment. User interactions may include, for example, without limitation, moving items of virtual content, resizing items of virtual content and / or locations of virtual private spaces, selecting input interface elements in a virtual user interface (e.g., a virtual representation of a mobile phone, a virtual keyboard, and / or other virtual interface), providing input through a virtual user interface, etc.
[0057] Operations for the XR engine 220, self-image engine 222, image processing engine 224, and rendering engine 226 may be implemented by any of the computing component 210. In one illustrative example, operations for the rendering engine 226 may be implemented by the GPU 214, and operations for the XR engine 220, self-image engine 222, and image processing engine 224 may be implemented by the CPU 212, the DSP 216, and / or the ISP 218. In some cases, the computing component 210 may include other electronic circuitry or hardware, computer software, firmware, or any combination thereof for performing any of the various operations described herein.
[0058] In some examples, the XR engine 220 can perform XR operations to generate an XR experience based on data from one or more sensors on the extended reality system 200, such as the image sensor 202, the accelerometer 204, the gyroscope 206, and / or one or more IMUs, radar, etc. In some examples, the XR engine 220 can perform tracking, localization, pose estimation, mapping, content anchoring operations, and / or any other XR operations / functions. The XR experience can include use of the extended reality system 200 to present XR content (e.g., virtual reality content, augmented reality content, mixed reality content, etc.) to a user during a virtual session. In some examples, the XR content and experiences can be provided by the extended reality system 200 through XR applications (e.g., executed or implemented by the XR engine 220) that provide a particular XR experience, such as, for example, an XR gaming experience, an XR classroom experience, an XR shopping experience, an XR entertainment experience, an XR activity (e.g., a movement, a troubleshooting activity, etc.), among others. During an XR experience, a user can view and / or interact with virtual content using extended reality system 200. In some cases, while the user can view and / or interact with the virtual content, the user can also view and / or interact with the physical environment around the user, allowing the user to have an immersive experience between the physical environment and the virtual content mixed or integrated with the physical environment.
[0059] The self-image engine 222 can perform various operations associated with capturing a self-image. In some cases, the self-image engine 222 can generate a self-image by combining a user's avatar (or another type of digital representation of the user) with one or more background frames. For example, the self-image engine 222 (in conjunction with one or more other components of the extended reality system 200) can perform a multi-action self-image capture process. One operation of the self-image capture process can involve capturing one or more frames of a real-world environment in which the extended reality system 200 is located. Another operation of the self-image capture process can involve generating an avatar that reflects the user's captured pose (e.g., facial expression, gesture, posture, location, etc.). A further operation of the self-image capture process can involve combining the avatar with one or more frames of the real-world environment. As described in more detail below, the self-image engine 222 can perform the operations of the multi-action self-image capture process in various orders and / or manners.
[0060] 3A and 3B are block diagrams illustrating examples of self-image capture systems 300(A) and 300(B), respectively. In some cases, self-image capture systems 300(A) and 300(B) may represent various exemplary implementations or operations of a single system or device (e.g., a single extended reality system or device) and various exemplary implementations of the self-image capture techniques described herein. For example, self-image capture systems 300(A) and 300(B) may correspond to various implementations or operations of self-image engine 222 of extended reality system 200. As illustrated, self-image capture systems 300(A) and 300(B) may include one or more of the same components. For example, self-image capture systems 300(A) and 300(B) may include one or more engines, including self-image initiation engine 302, avatar engine 304, background frame engine 306, and composition engine 308. The engines of the selfie image capture systems 300(A) and 300(B) may be configured to generate a selfie image frame 316. The selfie image frame 316 may include one or more background frames (e.g., background frame 314) overlaid with a digital representation of at least one user (e.g., avatar 318). For example, the selfie image frame 316 may correspond to a "selfie" photo or a "selfie" video.
[0061] In some cases, selfie image capture systems 300(A) and 300(B) may each be configured to perform a multi-operation process for selfie image capture. The following description provides a general description of various operations of the selfie image capture process performed by selfie image capture systems 300(A) and 300(B). A more detailed description of specific implementations corresponding to selfie image capture systems 300(A) and 300(B) is then provided, with specific reference to individual figures.
[0062] In some cases, the self-image initiation engine 302 can detect user input (e.g., user input 310) corresponding to the initiation of a self-image capture process. For example, the self-image initiation engine 302 can detect user input corresponding to activation of a “selfie mode” within an XR device (or other type of device) implementing the self-image capture systems 300(A) and 300(B). The user input 310 can include various types of user input, such as voice commands, touch input, and gesture input, among other types of input. In some cases, the self-image initiation engine 302 can detect the user input 310 while a user is wearing and / or using the XR device within a real-world environment. For example, the self-image initiation engine 302 can monitor one or more user interfaces of the XR device for user input 310 while the XR device is being used.
[0063] Based on detecting the user input 310, the selfie image capture system 300(A) and / or 300(B) can initiate the next action in the selfie image capture process. In one example, the avatar engine 304 can determine a user pose 312. The user pose 312 can include and / or correspond to a physical quality and / or characteristic of the user. For example, the user pose 312 can include one or more of the user's current facial expression, emotion, gesture (e.g., hand gesture), limb position, etc. Furthermore, the user pose 312 can include and / or correspond to the user's physical location (e.g., 3D location) within the real-world environment. The avatar engine 304 can determine the user pose 312 using various tracking and / or scanning techniques and / or algorithms. For example, the avatar engine 304 may determine the user pose 312 using one or more gaze tracking techniques, SLAM techniques, 6DOF positioning techniques, body tracking techniques, facial expression tracking techniques, computer vision techniques, any combination thereof, or other tracking and / or scanning techniques. In one example, the avatar engine 304 may determine the user pose 312 by applying one or more of such tracking and / or scanning techniques to image data captured by an inward-facing camera of an XR device (e.g., an HMD). In some cases, the inward-facing camera of the XR device may be able to capture the position of the user's face and / or body. For example, the field of view (FOV) of the inward-facing camera may correspond to less than the entire face and / or body of the user (e.g., due to camera placement and / or the XR device visually blocking the user's face). Thus, in some cases, the avatar engine 304 may determine (e.g., infer and / or estimate) the user pose 312 based on image data corresponding to a portion of the user.For example, as described in more detail below, the avatar engine 304 may determine the user pose 312 using a machine learning algorithm trained to determine a user pose based on image data corresponding to portions of the user. Additionally, in some examples, the avatar engine 304 may determine the user pose 312 using one or more outward-facing cameras of the XR device. For example, the avatar engine 304 may determine the user's facial expression based at least in part on image data captured by the inward-facing camera and determine the user's limb positions and / or hand gestures based at least in part on image data captured by the outward-facing camera. The avatar engine 304 may determine the user pose 312 using any combination of the inward-facing camera, the outward-facing camera, and / or additional cameras of the XR device.
[0064] In some cases, the avatar engine 304 can capture the user pose 312 based on additional user input. For example, the avatar engine 304 can detect user input that instructs the avatar engine 304 to capture the user's current pose. The additional user input can include various types of user input, such as voice commands, touch input, and gesture input, among other types of input. In one illustrative example, a user can provide input when they are satisfied with their current pose and / or location. Further, the user input can include input that instructs the avatar engine 304 to capture a single frame corresponding to the user pose 312 (e.g., to generate a single selfie image) or input that instructs the avatar engine 304 to capture a series of frames corresponding to the user pose 312 (e.g., to generate a selfie video).
[0065] In some examples, the avatar engine 304 can generate a user avatar 318 that reflects the user pose 312. As used herein, the term "avatar" can include any digital representation of all or part of a user. In one example, the user avatar can include computer-generated image data. Additionally or alternatively, the user avatar can include image data captured by an image sensor. Furthermore, the user avatar can correspond to an abstract (e.g., cartoon-like) representation of the user, or a photorealistic (or near-photorealistic) representation of the user. In some cases, generating the avatar 318 to reflect the user pose 312 can be referred to as "manipulating" the avatar 318.
[0066] In some examples, the avatar engine 304 may generate the avatar 318 using one or more machine learning systems and / or algorithms. For example, the avatar engine 304 may generate the avatar 318 based on a machine learning model trained using a machine learning algorithm on image data associated with various user poses. In this example, the machine learning model may be trained to output an avatar that corresponds to a captured user pose when information indicative of the captured user pose is input to the inferring model. In one illustrative example, once the machine learning model is trained, the avatar engine 304 may provide information indicative of the user's captured pose and one or more images of the user as input to the model. In some cases, the one or more images of the user may be unrelated to the user's pose. For example, the avatar engine 304 may capture one or more images of the user (e.g., a full-body image of the user) as part of setting up and / or configuring a selfie image capture system for the user. Based on the user's captured pose and the one or more images of the user, the machine learning model may output an avatar that resembles the user in the captured pose. For example, if a user's captured pose includes a particular hand gesture (e.g., a "peace sign"), the machine learning model may output an avatar that resembles the user making the particular hand gesture (even if the model has no previous image data associated with the user making the particular hand gesture).
[0067] The avatar engine 304 can implement various types of machine learning algorithms to generate the avatar 318. In one illustrative example, the avatar engine 304 can implement a deep neural network (NN), such as a generative adversarial network (GAN). Illustrative examples of deep neural networks are described below with respect to FIGS. 8 and 9. Additional examples of machine learning models include, without limitation, a time-delay neural network (TDNN), a deep feedforward neural network (DFFNN), a recurrent neural network (RNN), an autoencoder (AE), a variational adaptive ensemble (VAE), a denoising adaptive ensemble (DAE), a sparse adaptive ensemble (SAE), a Markov chain (MC), a perceptron, or some combination thereof. The machine learning algorithm may be a supervised learning algorithm, an unsupervised learning algorithm, a semi-supervised learning algorithm, any combination thereof, or other learning techniques.
[0068] In some examples, the avatar engine 304 may implement multiple (e.g., two or more) machine learning models configured to generate various different avatars 318. For example, as shown in FIGS. 3A and 3B , the avatar engine 304 may optionally include an avatar engine 304(A) and an avatar engine 304(B). In some cases, the avatar engine 304(A) may generate or obtain a first version of an avatar (denoted as avatar 318(A)) and a second version of an avatar (denoted as avatar 318(B)). In one example, the avatar 318(A) may correspond to a preview or initial version of the avatar 318. In one such example, the avatar 318(B) may correspond to a final version of the avatar 318. In some cases, the avatar 318(A) (preview or initial version) may be a lower-fidelity version compared to the avatar 318(B) (final version). In some examples, the avatar engine 304(A) may generate and / or display the avatar 318(A) before the user pose 312 and / or background frame 314 are captured (e.g., to facilitate capture of the desired user pose 312 and / or background frame 314). Once the user pose 312 and / or background frame 314 are captured, the avatar engine 304(B) may generate the avatar 318(B). The composition engine 308 may use the avatar 318(B) to generate the self-image frame 316. For example, a lower-fidelity avatar (e.g., avatar 318(A)) is displayed by the XR device or system when the user is in their current pose. The user may then operate the XR device or system to capture a final pose (e.g., used to generate the avatar 318(B)), which the composition engine 308 may use for composition when generating the self-image frame 316.The benefit of presenting a lower fidelity avatar (e.g., avatar 318(A)) during the pose capture stage or background image capture stage can allow composition to be performed with lower processing power before the final composition is performed using a higher fidelity avatar (e.g., avatar 318(B)).
[0069] In some cases, the avatar engine 304(A) may implement a first machine learning model that generates the avatar 318(A) (e.g., a preview or initial version of the avatar 318). In some cases, the avatar engine 304(B) may implement a second machine learning model that generates the avatar 318(B) (e.g., a final version of the avatar 318). In some aspects, the first machine learning model implemented by the avatar engine 304(A) may require less processing power than the second machine learning model implemented by the avatar engine 304(B). For example, the first machine learning model may be a relatively simple (e.g., low-fidelity) model that may be implemented locally (e.g., within an XR system or device). In some aspects, the first machine learning model may also be implemented in real-time (or near-real-time). The second machine learning model may be relatively complex (e.g., high-fidelity). In some aspects, the second machine learning model may be implemented offline (e.g., using a remote server or device configured to generate avatars).
[0070] In some cases, the background frame engine 306 can capture one or more background frames (e.g., background frame 314). The background frame 314 can include and / or correspond to any frame that serves as a background for a selfie image (or selfie video). In one example, the background frame engine 306 can capture the background frame 314 based on additional user input. For example, the background frame engine 306 can detect user input that instructs the background frame engine 306 to capture one or more frames of the real-world environment using an outward-facing camera of the XR system or XR device. The additional user input can include various types of user input, such as voice commands, touch input, gesture input, among other types of input. In one illustrative example, a user can provide input when satisfied with the current view of the real-world environment (which may be displayed on and / or via the display of the XR device). Further, the user input may include input that instructs the background frame engine 306 to capture a single frame of the real-world environment (e.g., to generate a single selfie image), or input that instructs the background frame engine 306 to capture a series of frames of the real-world environment (e.g., to generate a selfie video).
[0071] The composition engine 308 can generate a self-image frame 316 (or a series of self-image frames) based on the avatar 318 (e.g., avatar 318(B)) and the background frame 314. For example, the composition engine 308 can overlay the avatar 318 onto the background frame 314. As described above, the avatar engine 304 can determine a 3D location of the user that corresponds to the user pose 312. Accordingly, the composition engine 308 can overlay the avatar 318 within the background frame 314 in the corresponding location. For example, the composition engine 308 can represent the avatar 318 within the background frame 314 such that the avatar 318 appears to be located in the same 3D location as the user was when the avatar engine 304 captured the user pose 312. In this way, the resulting self-image frame 316 may appear to be an image of the user taken from a view facing the user (e.g., the view of a forward-facing camera used to capture a traditional “selfie”).
[0072] In some examples, the composition engine 308 can overlay the avatar 318 onto the background frame 314 using one or more machine learning systems and / or algorithms. For example, the composition engine 308 can overlay the avatar 318 onto the background frame 314 based on a machine learning model trained using a machine learning algorithm on image data associated with various avatars and / or background frames. In this example, the machine learning model can be trained to incorporate the avatar into the background frame such that the visual characteristics (e.g., lighting, color, etc.) of the avatar match and / or are consistent with the visual characteristics of the background frame. For example, once the machine learning model is trained, the composition engine 308 can provide one or more background frames and information associated with the avatar (e.g., an at least partially rendered avatar and / or a machine learning model trained to generate avatars) as input to the model. Based on the information associated with the avatar, the machine learning model can overlay the avatar onto one or more background frames in a natural and / or realistic manner. In one illustrative example, the model can determine that one or more background frames include a dark area in the location where the avatar will be overlaid. In this example, the model can represent the avatar to include corresponding dark areas.
[0073] The composition engine 308 may implement various types of machine learning algorithms for generating the selfie-picture frame 316, including any of the machine learning algorithms that may be implemented by the avatar engine 304 to generate the avatar 318 (described above). In some cases, the machine learning model implemented by the composition engine 308 may differ from the machine learning model implemented by the avatar engine 304. For example, the output of a machine learning model trained to generate the avatar 318 may be input to a machine learning model trained to generate the selfie-picture frame 316.
[0074] In one illustrative example, the self-image capture system 300(A) shown in FIG. 3A may perform an operation to generate an avatar 318 before an operation to capture a background frame 314. FIGS. 4A, 4B, and 4C show an example of such a self-image capture process. In this example, the self-image initiation engine 302 may detect a user input 310 corresponding to the initiation of a self-image capture mode of the XR device. In response to the user input 310, the self-image capture system 300(A) may initiate a pose capture mode in which the avatar engine 304 may capture a user pose 312. In one illustrative example, the avatar engine 304 may output instructions (e.g., in a display of the XR device) instructing the user to assume a desired pose for the self-image. However, in some examples, the avatar engine 304 may not output instructions (e.g., if the user is familiar with the self-image capture process). In some cases, the desired pose may include a desired 3D location in the real-world environment (e.g., to facilitate representing the avatar 318 in a corresponding location in the background frame 314). While in pose capture mode, the avatar engine 304 may detect user input that instructs the avatar engine 304 to capture a user pose 312.
[0075] FIG. 4A shows an example frame 402 corresponding to at least a portion of a user pose 312. In this example, the user pose 312 includes a hand gesture (e.g., a "peace sign"). The avatar engine 304 can detect the hand gesture based at least in part on image data captured by one or more outward-facing cameras of the XR device. Although not shown, the user pose 312 may include additional information about the user's physical appearance. For example, the avatar engine 304 can determine information about the position of the user's body and / or other limbs. In another example, the avatar engine 304 can determine information about the user's facial expressions (e.g., based on image data captured by one or more inward-facing cameras of the XR device). Based on the user pose 312, the avatar engine 304 can generate (e.g., manipulate) an avatar 318. For example, the avatar engine 304 can provide the user pose 312 (and one or more images of the user) to a machine learning model trained to generate an avatar that reflects the user pose. In one illustrative example, avatar engine 304(A) can generate avatar 318(A) (eg, a preview version of avatar 318) based on user pose 312.
[0076] Once the avatar engine 304 generates an avatar 318 (e.g., avatar 318(A)), the selfie image capture system 300(A) can initiate a background capture mode in which the background frame engine 306 can capture a background frame 314. In one illustrative example, the background frame engine 306 can output instructions (e.g., in the display of the XR device) instructing the user to select a 3D location in the real-world environment for the selfie image (or selfie video). However, in some examples, the background frame engine 306 may not output instructions (e.g., if the user is familiar with the selfie image capture process). In one example, the user can move within the real-world environment until the current FOV of the display of the XR device corresponds to the desired background frame for the selfie image frame 316. While in the background capture mode, the background frame engine 306 can detect user input that instructs the background frame engine 306 to capture the desired background frame. Once the background frame engine 306 captures the background frame 314, the composition engine 308 may generate the self-image frame 316 by overlaying the avatar 318 onto the background frame 314. For example, the composition engine 308 may represent the avatar 318 within the background frame 314 in a location corresponding to the 3D location of the user pose 312. In one illustrative example, the avatar engine 304(B) may generate and / or obtain an avatar 318(B) (e.g., a final version of the avatar 318) based on the captured background frame 314. In this example, the composition engine 308 may generate the self-image frame 316 by representing the avatar 318(B) within the background frame 314.
[0077] In some cases, the avatar engine 304 can represent an avatar 318 within the XR device's display while the selfie image capture system 300(A) is operating in a background capture mode (e.g., while the user is moving around the real-world environment to select a background frame 314). For example, the avatar engine 304 can represent the avatar 318 (e.g., avatar 318(A)) in a location corresponding to the user's 3D location when the user pose 312 was captured. As the FOV of the XR device's display changes (e.g., based on the user's movement), the avatar engine 304 can adjust the location of the avatar 318 (e.g., avatar 318(A)) represented within the display to accommodate the movement. Thus, in the selfie image capture process implemented by the selfie image capture system 300(A), the avatar 318 can be "world-locked." In some cases, the world-locked avatar can enable the user to select a background frame appropriate for the real-world location corresponding to the avatar 318. For example, if the XR device moves such that the 3D location corresponding to the avatar 318 is no longer within the XR device's FOV, the avatar engine 304 can remove the avatar 318 from the display. In this way, the user can ensure that the avatar 318 is properly positioned within the FOV corresponding to the background frame 314.
[0078] 4B shows an example frame 404 corresponding to the FOV of the XR device while the self-image capture system 300(A) is operating in background capture mode. FIG. 4C shows an example frame 406 corresponding to a self-image frame 316 generated by the composition engine 308 when the background frame engine 306 captured a background frame 314. In these examples, frame 404 includes an avatar 318(A) and frame 406 includes an avatar 318(B). For example, the avatar engine 304 can represent the avatar 318(A) while the self-image capture system 300(A) is in background capture mode, and the composition engine 308 can replace the avatar 318(A) with the avatar 318(B) when generating the self-image frame 316. As noted above, the avatar 318(A) can be a low-fidelity version of the avatar 318, and the avatar 318(B) can be a high-fidelity version of the avatar 318. For example, the avatar engine 304 may generate the avatars 318(A) and 318(B) using different machine learning models implemented by the avatar engines 304(A) and 304(B). In one example, the avatar engine 304(A) may generate the avatar 318(A) using a low-fidelity machine learning model that involves and / or requires a smaller amount of processing power and / or time than a high-fidelity machine learning model used by the avatar engine 304(B) to generate the avatar 318(B). Using a low-fidelity machine learning model to generate the avatar 318(A) may enable the avatar engine 304(A) to update the location of the avatar 318(A) within the display of the XR device in real time (near real time) as the FOV of the XR device changes during background capture mode. Furthermore, using a high-fidelity machine learning model to generate the avatar 318(B) may create a higher quality (e.g., more realistic) avatar for the final selfie image. For example, avatar 318(A) in FIG. 4B is cartoon-like, while avatar 318(B) in FIG. 4C is photorealistic.In one example, avatar engine 304(A) can implement a low-fidelity machine learning model (e.g., locally) on the XR device, while avatar engine 304(B) can instruct a remote server configured to generate avatars to implement a high-fidelity machine learning model. Avatar engine 304 can generate any type or number of avatars (including a single avatar using a single machine learning model).
[0079] Referring to FIG. 3B , the self-image capture system 300(B) may perform an operation to capture a background frame 314 before an operation to generate an avatar 318. FIGS. 4D, 4E, and 4F illustrate an example of such a self-image capture process. In this example, the self-image initiation engine 302 may detect a user input 310 corresponding to the initiation of a self-image capture mode of the XR device. In response to the user input 310, the self-image capture system 300(B) may initiate a background capture mode in which the background frame engine 306 may capture a background frame 314. In this background capture mode, the background frame engine 306 may optionally output instructions prompting the user to select a 3D location within the real-world environment for the self-image (or self-video). In one example, the user may move within the real-world environment until the current FOV of the XR device's display corresponds to the desired background frame for the self-image frame 316. While in the background capture mode, the background frame engine 306 may detect a user input instructing the background frame engine 306 to capture the desired background frame. 4D shows an example frame 408 corresponding to the captured background frame. The background capture mode of selfie image capture system 300(B) may be similar to the background capture mode of selfie image capture system 300(A), but this background capture mode may differ in that the avatar 318 (e.g., avatar 318(A)) is not displayed within (e.g., manipulated on) the display of the XR device while the user selects a location for the selfie image.
[0080] Once the background frame engine 306 captures the background frame 314, the selfie image capture system 300(B) may enter a pose capture mode in which the avatar engine 304 may capture a user pose 312. In some cases, the avatar engine 304 may optionally output instructions instructing the user to assume a desired pose for the selfie image. In some examples, the desired pose may include a desired 3D location in the real-world environment (e.g., to facilitate representing the avatar 318 in a corresponding location in the background frame 314). While in the pose capture mode, the avatar engine 304 may detect user input that instructs the avatar engine 304 to capture a user pose 312. Based on the user pose 312, the avatar engine 304 may generate (e.g., manipulate) the avatar 318. For example, the avatar engine 304 may provide the user pose 312 (and one or more images of the user) to a machine learning model trained to generate an avatar that reflects the user pose. In one illustrative example, avatar engine 304(B) may generate avatar 318(B) based on user pose 312. Once avatar engine 304 generates avatar 318 (e.g., avatar 318(B)), composition engine 308 may generate self-image frame 316 by overlaying avatar 318 onto background frame 314. For example, composition engine 308 may represent avatar 318 within background frame 314 in a location corresponding to the 3D location of user pose 312.
[0081] In some cases, the avatar engine 304(A) may represent the avatar 318(A) within the display of the XR device while the selfie image capture system 300(B) is in pose capture mode (e.g., while the user moves around the real-world environment before the user pose 312 is captured). For example, the avatar engine 304 may dynamically update the avatar 318(A) (e.g., in real time or near real time) based on changes in the user's pose. Changes in the user's pose may include changes in the user's 3D location within the real-world environment and / or changes in the user's physical appearance (e.g., the user's facial expression, hand gestures, limb positions, etc.). The avatar engine 304(A) may display the avatar 318(A) within a preview window that displays the background frame 314. For example, the avatar engine 304(A) may display a preview window within the display of the XR device to update the avatar 318(A) as the user moves around the real-world environment. In some cases, this version of avatar 318(A) may be "head-locked" (e.g., as opposed to a "world-locked" version of avatar 318(A) generated by selfie image capture system 300(A)). In one example, a head-locked avatar can facilitate accurate selfie image composition by allowing a user to select a user pose appropriate for a previously selected background frame 314.
[0082] 4E shows an example frame 410 that includes a preview window 414. The preview window 414 can display a (static) background frame 314 and a dynamically updated avatar 318(A). For example, as a user's pose changes (e.g., due to the user's movement within the real-world environment), the avatar engine 304(A) can update the avatar 318(A) in the preview window 414 to address the change. While displaying the avatar 318(A), the avatar engine 304(B) can detect user input that instructs the avatar engine 304(B) to capture a current user pose (e.g., corresponding to the user pose 312). Based on receiving such user input, the avatar engine 304(B) can generate (e.g., manipulate) the avatar 318(B) based on the user pose 312. 4F shows an example frame 412 corresponding to the self-image frame 316 generated by the composition engine 308 when the avatar engine 304(B) generated the avatar 318(B) based on the user pose 312. As described above, the avatar 318(A) may correspond to a low-fidelity version of the avatar 318, and the avatar 318(B) may correspond to a high-fidelity version of the avatar 318. For example, the avatar engine 304(A) may generate the avatar 318(A) using a local and / or low-fidelity machine learning algorithm, and the avatar engine 304(B) may generate and / or acquire the avatar 318(B) using a remote and / or high-fidelity machine learning algorithm. The self-image capture system 300(B) may generate any type or number of avatars using any suitable machine learning algorithm.
[0083] As described above, the self-image capture systems 300(A) and 300(B) can implement various sets of processes for capturing self-images within an XR environment. These self-image capture systems can enable HMDs and other devices without mechanisms designed for self-image capture (such as a forward-facing camera on a mobile phone) to generate natural and realistic self-images. Furthermore, by capturing a user's pose and background frames at different points in time, the self-image capture systems of the present disclosure can enable a user to precisely customize and / or optimize the composition of their self-image.
[0084] As mentioned above, in some cases, the self-image capture techniques and systems of the present disclosure may be used to generate self-videos. Both the self-image capture system 300(A) and the self-image capture system 300(B) may be used to generate self-videos. FIG. 5A is a flowchart of an example process 500(A) for self-video capture that may be performed by the self-image capture system 300(A). At operation 502, the process 500(A) may include self-video initiation. For example, the self-image capture system 300(A) may detect a user input corresponding to the start of a self-video mode. At operation 504, the process 500(A) may include starting to record a user's pose. For example, the self-image capture system 300(A) may capture one or more frames using an inward-facing camera and / or an outward-facing camera of an XR device. The self-image capture system 300(A) may generate (e.g., manipulate) multiple avatars (e.g., a series of avatars) corresponding to the user's pose indicated by all or some of the captured frames. At operation 506, the selfie image capture system 300(A) may stop recording the user's pose. The selfie image capture system 300(A) may record any number of frames associated with the user's pose between operation 504 and operation 506. In one example, the selfie image capture system 300(A) may record a predetermined number of frames (e.g., 10 frames, 50 frames, etc.) and / or record frames for a predetermined amount of time (e.g., 2 seconds, 5 seconds, etc.). In another example, the selfie image capture system 300(A) may record frames until it detects a user input that instructs the selfie image capture system 300(A) to stop recording.
[0085] At operation 508, the process 500(A) may include starting recording one or more background frames (e.g., using a front-facing camera of the XR device). At operation 510, the selfie image capture system 300(A) may stop recording the background frames. The selfie image capture system 300(A) may record any number of background frames between operations 508 and 510. In one example, the selfie image capture system 300(A) may record a number of frames corresponding to the number of recorded frames associated with the user's pose (e.g., the number of frames recorded in operation 504). For example, the recording process for recording background frames may automatically terminate (e.g., time out) after an amount of time corresponding to recording frames associated with the user's pose. In another example, the selfie image capture system 300(A) may record a single background frame. In this example, the single background frame may correspond to a static background for the selfie-video. At operation 512, the process 500(A) may include self-video composition. For example, the selfie image capture system 300(A) can overlay multiple avatars onto one or more background frames.
[0086] 5B is a flowchart of an example process 500(B) for self-video capture that may be performed by the self-image capture system 300(B). At operation 514, the process 500(B) may include self-video initiation. For example, the self-image capture system 300(B) may detect a user input corresponding to the start of a self-video mode. At operation 516, the process 500(B) may include starting to record one or more background frames (e.g., using a front-facing camera of an XR device). At operation 518, the self-image capture system 300(B) may stop recording the background frames. The self-image capture system 300(B) may record any number of background frames between operations 516 and 518. In one example, the self-image capture system 300(B) may record a predetermined number of background frames and / or record background frames for a predetermined amount of time. In another example, the selfie image capture system 300(B) may record background frames until it detects a user input that instructs the selfie image capture system 300(A) to stop recording.
[0087] At operation 520, the process 500(B) may include starting to record a user's pose. For example, the selfie image capture system 300(B) may capture one or more frames using an inward-facing camera and / or an outward-facing camera of the XR device. The selfie image capture system 300(B) may generate (e.g., manipulate) multiple avatars corresponding to the user's pose indicated by all or a portion of the captured frames. At operation 522, the selfie image capture system 300(B) may stop recording frames associated with the user's pose. The selfie image capture system 300(B) may record any number of frames associated with the user's pose. In one example, the selfie image capture system 300(B) may record a number of frames corresponding to the number of recorded background frames. For example, the recording process for recording the user's pose may automatically terminate (e.g., time out) after an amount of time corresponding to recording the background frames. At operation 524, the process 500(B) may include self-video composition. For example, the selfie image capture system 300(B) can overlay multiple avatars onto one or more background frames.
[0088] In some cases, the disclosed techniques and systems for self-image capture in an XR environment can be used to generate self-images or self-videos including multiple people. Both the self-image capture system 300(A) and the self-image capture system 300(B) can be used to generate self-images or self-videos including multiple people. FIG. 6A is a flowchart of an example process 600(A) for multi-user self-image capture that can be performed by the self-image capture system 300(A). At operation 602, the process 600(A) can include multi-user self-image initiation. For example, the self-image capture system 300(A) can detect a user input corresponding to the initiation of a multi-user self-image capture mode. At operation 604, the process 600(A) can include generating an avatar based on the captured user pose. For example, the self-image capture system 300(A) can manipulate an avatar that corresponds to the current pose of a user wearing an XR device. In some cases, the selfie image capture system 300(A) may capture a pose of a user while the user poses with one or more other people to be included in a multi-user selfie image.
[0089] At operation 606, the process 600(A) may include obtaining avatar data corresponding to the additional person. For example, the selfie image capture system 300(A) (implemented on an XR device worn by a user) may send a request to one or more nearby XR devices to receive data associated with avatars of any other people to be included in the multi-user selfie image. In one example, the selfie image capture system 300(A) may broadcast the request to any XR devices within broadcast range of the selfie image capture system 300(A). In another example, the selfie image capture system 300(A) may send a specific request to XR devices known to be associated with one or more people to be included in the multi-user selfie image. For example, the selfie image capture system 300(A) may send the request to a specific XR device based on user input and / or may send the request to an XR device with which the selfie image capture system 300(A) previously communicated and / or connected.
[0090] In one example, the avatar data requested by the selfie image capture system 300(A) may include avatars corresponding to one or more other people to be included in the multi-user selfie image. For example, the selfie image capture system 300(A) may prompt a selfie image capture system implemented on an XR device associated with the one or more other people to generate an avatar corresponding to the captured pose of the one or more other people. In another example, the avatar data may include data that enables the selfie image capture system 300(A) to generate an avatar corresponding to the one or more other people. For example, the avatar data may include information about the captured pose of the one or more other people. The avatar data may also include a machine learning model (e.g., an avatar tuning network) trained to generate avatars for the one or more other people. In some cases, the model may be trained using one or more images (e.g., full-body images) of the one or more other people. Notably, in some cases, the person (or people) to be included in the multi-user selfie image may not be wearing and / or associated with an XR device configured to generate the avatar. In these cases, the selfie image capture system 300(A) may not obtain data associated with an avatar corresponding to the person.
[0091] At operation 608, the process 600(A) may include capturing a background frame. For example, the selfie image capture system 300(A) may detect user input that instructs the selfie image capture system 300(A) to capture a background frame that corresponds to the current FOV of the XR device. In one example, the background frame may include image data corresponding to one or more other people to be included in the multi-user selfie image. For example, a user may move within a real-world environment while one or more other people remain stationary. Once the user determines that the current FOV of the XR device is appropriate for a background frame (e.g., based on the current FOV including one or more other people), the user may provide input that instructs the selfie image capture system 300(A) to capture the background frame.
[0092] At operation 610, the process 600(A) can include multi-user selfie-image composition. In an example where the selfie-image capture system 300(A) receives avatars of one or more people (e.g., previously generated avatars), the selfie-image capture system 300(A) can overlay the avatars (and the user's avatar) on a background frame. For example, the selfie-image capture system 300(A) can replace image data corresponding to one or more other people with appropriate avatars. In an example where the selfie-image capture system 300(A) receives a machine learning model trained to generate avatars of one or more other people, the selfie-image capture system 300(A) can use the model to represent avatars corresponding to one or more other people in the background frame.
[0093] For example, FIG. 6C is a block diagram of a multi-user selfie image capture system 622. The multi-user selfie image capture system 622 can receive user poses 630(1)-(N), which correspond to captured user poses of one or more other people. The multi-user selfie image capture system 622 can also receive avatar networks 624(1)-624(N), which correspond to machine learning models (e.g., model files of machine learning models) trained to generate avatars for one or more other people. Based on the user poses 630(1)-(N), the multi-user selfie image capture system 622 can implement the avatar networks 624(1)-624(N) to generate avatars corresponding to one or more other people. The generated avatars can be input to a selfie image generator 626, which corresponds to a machine learning model trained to represent one or more avatars within a background frame. The self-image generator 626 can generate a multi-user self-image 628 that includes avatars corresponding to the user and one or more other people. In some cases, the self-image generator 626 can ensure that the avatars are globally matched and / or consistent within the multi-user self-image 628. For example, the self-image generator 626 can standardize the lighting, color, and / or other visual characteristics of the avatars. In some cases, the self-image generator 626 can also remove any occlusions visible in the background frame that may obscure portions of the avatars. Additionally, the self-image generator 626 can ensure that XR devices worn by one or more people are not depicted in the multi-user self-image 628. In examples where a person (or people) to be included in the multi-user self-image 628 is not associated with an XR device and / or avatar, the multi-user self-image capture system 622 can leave image data corresponding to the person in the background frame unchanged.
[0094] FIG. 6B is a flowchart of an example process 600(B) for multi-user selfie image capture that may be performed by the selfie image capture system 300(B). At operation 612, the process 600(B) may include multi-user selfie image initiation. For example, the selfie image capture system 300(B) may detect a user input corresponding to initiating a multi-user selfie image capture mode. At operation 614, the process 600(B) may include capturing a background frame. For example, the selfie image capture system 300(B) may detect a user input that instructs the selfie image capture system 300(B) to capture a background frame corresponding to the current FOV of the XR device. In some cases, this background frame does not include image data corresponding to one or more people to be included in the multi-user selfie image. For example, the selfie image capture system 300(B) may capture a background frame before the user and one or more other people pose themselves as desired for the multi-user selfie image. At operation 616, process 600(B) may include generating an avatar based on the captured user pose. For example, selfie image capture system 300(B) may detect a user input that directs selfie image capture system 300(B) to capture a pose of the user (e.g., as the user and one or more other people pose themselves). Selfie image capture system 300(B) may manipulate an avatar that corresponds to the current pose of the user wearing the XR device.
[0095] At operation 618, process 600(B) may include obtaining avatar data corresponding to one or more other people. For example, selfie image capture system 300(B) (implemented on the user's XR device) may send a request to one or more nearby XR devices to receive avatars of one or more other people. In another example, selfie image capture system 300(B) may send a request to receive data that enables selfie image capture system 300(B) to generate avatars of one or more other people. For example, selfie image capture system 300(B) may send a request to receive a machine learning model (such as avatar network 624(1)-(N) shown in FIG. 6C ) that has been trained to generate avatars of one or more other people. In some cases, selfie image capture system 300(B) may send a request to nearby XR devices in any of the same manners that may be performed by selfie image capture system 300(A) at operation 606 of process 600(A). At operation 620, the process 600(B) can perform multi-user selfie-image composition. In an example where the selfie-image capture system 300(B) receives avatars of one or more people (e.g., previously generated avatars), the selfie-image capture system 300(B) can overlay the avatars (and the user's avatar) on a background frame. In an example where the selfie-image capture system 300(B) receives a machine learning model trained to generate avatars of one or more other people, the selfie-image capture system 300(B) can use the model to represent avatars corresponding to the one or more other people in the background frame. For example, the selfie-image capture system 300(B) can input the avatars generated by the model to a selfie-image generator 626 shown in FIG. 6C .
[0096] 7 is a flow diagram illustrating an exemplary process 700 for self-image capture within an extended reality environment. For clarity, process 700 is described with reference to self-image capture systems 300(A) and 300(B) shown in FIGS. 3A and 3B. The steps or operations outlined herein are examples and can be implemented in any combination thereof, including combinations in which some steps or operations are removed, added, or modified.
[0097] At operation 702, process 700 includes capturing a pose of a user of the extended reality system, the user's pose including the user's location within a real-world environment associated with the extended reality system. In some examples, the avatar engine 304 can capture the user's pose based at least in part on image data captured by an inward-facing camera system of the extended reality system. Further, the avatar engine 304 can capture the user's pose based at least in part on determining the user's facial expressions and / or determining the user's gestures. In one example, the avatar engine 304 can determine the user's location within the real-world environment based at least in part on generating a three-dimensional map of the real-world environment.
[0098] At operation 704, process 700 includes generating a digital representation of the user, where the digital representation of the user reflects a pose of the user. In some examples, process 700 can generate a first digital representation of the user and a second digital representation of the user. In some cases, the second digital representation of the user can be a digital representation of the user with higher fidelity than the first digital representation of the user. For example, process 700 can include generating or obtaining a first digital representation having a first fidelity and generating or obtaining a second digital representation having a second fidelity (the second fidelity being higher than the first fidelity). In some aspects, the first digital representation of the user can correspond to a preview digital representation of the user that can be displayed to the user to facilitate capturing a desired background frame and / or pose of the user. In some aspects, the second digital representation of the user can correspond to a final digital representation of the user.
[0099] In some examples, the avatar engine 304 may generate the digital representation of the user using a machine learning algorithm. For example, in some cases, the avatar engine 304 may generate the first digital representation of the user based on a first machine learning algorithm. The avatar engine 304 may obtain the second digital representation of the user based on a second machine learning algorithm. In some cases, the avatar engine 304 may generate the first digital representation of the user based on implementing the first machine learning algorithm on the extended reality system. In some cases, the avatar engine 304 may obtain the second digital representation of the user by causing a server configured to generate digital representations of users to generate the second digital representation of the user using a second machine learning algorithm.
[0100] At operation 706, process 700 includes capturing one or more frames of the real-world environment. In some examples, the background frame engine 306 may capture one or more frames of the real-world environment using an outward-facing camera system of the extended reality system. In one example, operation 706 may be performed after operation 702 and / or operation 704. For example, the avatar engine 304 may generate a digital representation of the user before the background frame engine 306 captures one or more frames of the real-world environment. In this example, the avatar engine 304 may display the digital representation of the user in a display location corresponding to the user's location in the real-world environment within a display of the extended reality system through which the real-world environment is viewed. In some examples, the avatar engine 304 may display the digital representation of the user using the user's captured pose (captured in operation 702). While the digital representation of the user is displayed within the display of the extended reality system, the background frame engine 306 may detect user input corresponding to an instruction to capture one or more frames of the real-world environment. The background frame engine 306 can then capture one or more frames of the real-world environment based on the user input. In one example, the avatar engine 304 can display a first (e.g., preview) digital representation of the user in a display of the extended reality system before the background frame engine 306 captures one or more frames of the real-world environment. The avatar engine 304 can generate a second (e.g., final) digital representation of the user based on the one or more frames of the real-world environment being captured and / or based on the user's pose in the one or more frames.
[0101] In another example, operation 706 may be performed before operation 702 and / or operation 704. For example, the background frame engine 306 may capture one or more frames of the real-world environment before the avatar engine 304 captures the user's pose. In this example, the avatar engine 304 may display the digital representation of the user in a display location corresponding to the user's location in the real-world environment within the display of the extended reality system on which the one or more frames of the real-world environment are displayed. In some examples, the avatar engine 304 may display the digital representation of the user using the current user's pose. The avatar engine 304 may update the display location of the digital representation of the user based on detecting a change in the user's location within the real-world environment. While the digital representation of the user is displayed within the display of the extended reality system, the avatar engine 304 may detect a user input corresponding to an instruction to capture the user's pose. The avatar engine 304 may then capture the user's pose based on the user input. In one example, the avatar engine 304 may display a first (e.g., preview) digital representation of the user within a display of the extended reality system before capturing the user's pose. The avatar engine 304 may generate a second (e.g., final) digital representation of the user based on one or more captured frames of the real-world environment and / or based on the user's pose being captured.
[0102] At operation 708, process 700 includes overlaying the digital representation of the user onto one or more frames of the real-world environment. In some examples, the configuration engine 308 may overlay the digital representation of the user onto one or more frames of the real-world environment in a frame location that corresponds to the user's location within the real-world environment. In one example, the configuration engine 308 may overlay the digital representation of the user onto one or more frames of the real-world environment using a machine learning algorithm. The machine learning algorithm may be different from the machine learning algorithm used by the avatar engine 304 to generate the digital representation of the user.
[0103] In some examples, process 700 can include capturing a pose of a person in a real-world environment and generating a digital representation of the person. The digital representation of the person can reflect the person's pose. Process 700 can also include overlaying the digital representation of the user and the digital representation of the person onto one or more frames of the real-world environment. In one example, avatar engine 304 can generate the digital representation of the person based at least in part on information associated with the digital representation of the person received from the person's extended reality system. The information associated with the digital representation of the person can include a machine learning model trained to generate the digital representation of the person.
[0104] In a further example, process 700 may include capturing a plurality of poses of the user associated with the plurality of frames and generating a plurality of digital representations of the user corresponding to the plurality of frames. Process 700 may also include overlaying the plurality of digital representations of the user onto one or more frames of a real-world environment, where the one or more frames of the real-world environment include the plurality of frames of the real-world environment.
[0105] In some examples, processes 500(A), 500(B), 600(A), 600(B), 700, and / or other processes described herein may be performed by one or more computing devices or apparatuses. In some examples, processes 500(A), 500(B), 600(A), 600(B), 700, and / or other processes described herein may be performed by one or more computing devices, such as the extended reality system 200 shown in FIG. 2, the selfie image capture system 300(A) shown in FIG. 3A, the selfie image capture system 300(B) shown in FIG. 3B, the multi-user selfie image capture system 622 shown in FIG. 6C, and / or the computing device architecture 1000 shown in FIG. 10. In some cases, such computing devices or apparatuses may include a processor, microprocessor, microcomputer, or other component of a device configured to perform the steps of processes 500(A), 500(B), 600(A), 600(B), 700. In some examples, such a computing device or apparatus may include one or more sensors configured to capture image data. For example, the computing device may include a smartphone, a camera, a head-mounted display, a mobile device, or other suitable device. In some examples, such a computing device or apparatus may include a camera configured to capture one or more images or videos. In some cases, such a computing device may include a display for displaying the images. In some examples, the one or more sensors and / or camera are separate from the computing device, in which case the computing device receives the sensed data. Such a computing device may further include a network interface configured to communicate data.
[0106] Components of a computing device may be implemented in circuit configurations. For example, components may include and / or be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, a graphics processing unit (GPU), a digital signal processor (DSP), a central processing unit (CPU), and / or other suitable electronic circuitry), and / or may include and / or be implemented using computer software, firmware, or any combination thereof to perform various operations described herein. A computing device may further include a display (as an example of an output device or in addition to an output device), a network interface configured to communicate and / or receive data, any combination thereof, and / or other components. The network interface may be configured to communicate and / or receive Internet Protocol (IP)-based data or other types of data.
[0107] Processes 500(A), 500(B), 600(A), 600(B), 700 are illustrated as logical flow diagrams, whose operations represent sequences of actions that may be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, the actions represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the described actions. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform particular functions or implement particular data types. The order in which the actions are described is not intended to be construed as a limitation, and any number of the described actions may be combined in any order and / or in parallel to implement a process.
[0108] Additionally, processes 500(A), 500(B), 600(A), 600(B), 700, and / or other processes described herein may be executed under the control of one or more computer systems configured with executable instructions and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that collectively execute on one or more processors, by hardware, or a combination thereof. As noted above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.
[0109] FIG. 8 is an illustrative example of a deep learning neural network 800 that may be used by a light estimator. The input layer 820 includes input data. In one illustrative example, the input layer 820 may include data representing pixels of an input frame. The neural network 800 includes multiple hidden layers 822a, 822b, through 822n. The hidden layers 822a, 822b, through 822n include “n” hidden layers, where “n” is an integer greater than or equal to 1. The number of hidden layers may be as many as required for a given application. The neural network 800 further includes an output layer 824 that provides output resulting from processing performed by the hidden layers 822a, 822b, through 822n. In one illustrative example, the output layer 824 may provide a light estimate associated with the light of the frame. The light estimate may include illumination parameters and / or latent feature vectors.
[0110] Neural network 800 is a multi-layered neural network of interconnected nodes. Each node can represent information. The information associated with a node is shared among different layers, with each layer retaining the information as it is processed. In some cases, neural network 800 can include a feedforward network, in which there are no feedback connections where the output of the network is fed back to itself. In some cases, neural network 800 can include a recurrent neural network, which can have loops that allow information to be carried across nodes while reading at the input.
[0111] Information can be exchanged between nodes through node-to-node interconnections between various layers. Nodes in the input layer 820 can activate a set of nodes in the first hidden layer 822a. For example, as shown, each input node in the input layer 820 is connected to each node in the first hidden layer 822a. Nodes in the hidden layers 822a, 822b, through 822n can transform the information of each input node by applying an activation function to that information. The information derived from the transformation can then be passed to and activated by nodes in the next hidden layer 822b, which can perform a specific, specified function. Exemplary functions include convolution, upsampling, data transformation, and / or any other suitable function. The output of the hidden layer 822b can then activate nodes in the next hidden layer, and so on. The output of the last hidden layer 822n can activate one or more nodes in the output layer 824 to which the output is provided. In some cases, a node in neural network 800 (e.g., node 826) is shown as having multiple output lines, whereas the node has a single output, and all lines shown as outputting from the node represent the same output value.
[0112] In some cases, each node or the interconnections between nodes may have weights, which are sets of parameters derived from training the neural network 800. Once the neural network 800 is trained, it may be referred to as a trained neural network and may be used to classify one or more objects. For example, the interconnections between nodes may represent information learned about the interconnected nodes. The interconnections may have adjustable numerical weights that can be adjusted (e.g., based on a training data set), allowing the neural network 800 to be adaptive to inputs and to learn as more data is processed.
[0113] Neural network 800 is pre-trained to process features from data in input layer 820 using different hidden layers 822a, 822b, through 822n to provide output through output layer 824. In examples where neural network 800 is used to identify objects in images, neural network 800 may be trained using training data that includes both images and labels. For example, training images may be input into the network, with each training image having a label that indicates the class of one or more objects in each image (essentially telling the network what the object is and what characteristics it has). In one illustrative example, the training images may include images of the number 2, in which case the label for the image may be [0 0 1 0 0 0 0 0 0 0].
[0114] In some cases, neural network 800 may adjust node weights using a training process called backpropagation. Backpropagation may include a forward pass, a loss function, a backward pass, and a weight update. The forward pass, loss function, backward pass, and parameter update are performed during one training iteration. The process may be repeated for a number of iterations for each set of training images until neural network 800 is well enough trained that the layer weights are precisely tuned.
[0115] For the example of identifying objects in an image, the forward pass may include passing a training image through neural network 800. Before neural network 800 is trained, the weights are first randomized. The image may include, for example, an array of numbers representing pixels of the image. Each number in the array may include a value between 0 and 255 representing the pixel intensity at that location in the array. In one example, the array may include a 28x28x3 array of numbers, with 28 rows and 28 columns of pixels and three color components (such as a red component, a green component, and a blue component, or a luma component and two chroma components).
[0116] For the first training iteration for the neural network 800, due to the weights being chosen randomly at initialization, the output may contain values that do not give preference to any particular class. For example, if the output is a vector with the probabilities that an object contains different classes, the probability values for each of the different classes may be equal or at least very similar (e.g., for 10 possible classes, each class may have a probability value of 0.1). With the initial weights, the neural network 800 is unable to determine low-level features and therefore is unable to make an accurate determination of what the classification of an object is likely to be. A loss function can be used to analyze errors in the output. Any suitable definition of the loss function can be used. An example of a loss function includes the mean squared error (MSE). MSE is
number
[0117] The loss (or error) is large for the first training images because the actual values are significantly different from the predicted outputs. The goal of training is to minimize the amount of loss so that the predicted outputs are the same as the training labels. The neural network 800 can perform a backward pass by determining which inputs (weights) contributed most to the network's loss, and then adjust the weights so that the loss becomes smaller and is eventually minimized.
[0118] To determine the weights that contributed most to the network's loss, the derivative of the loss with respect to the weights (denoted as dL / dW, where W is the weight at a particular layer) can be calculated. After the derivatives are calculated, a weight update can be performed by updating all of the weights of the filter. For example, the weights can be updated such that they change in the opposite direction of the gradient. The weight update can be
number
[0119] Neural network 800 can include any suitable deep network. One example includes a convolutional neural network (CNN), which includes an input layer and an output layer with multiple hidden layers between the input and output layers. An example of a CNN is described below with respect to FIG. 8. The hidden layers of a CNN include a series of convolutional layers, nonlinear layers, pooling (for downsampling) layers, and fully connected layers. Neural network 800 may include any other deep network other than a CNN, such as an autoencoder, a deep belief net (DBN), a recurrent neural network (RNN), among others.
[0120] FIG. 9 is an illustrative example of a convolutional neural network 900 (CNN 900). The input layer 920 of the CNN 900 includes data representing an image. For example, the data can include an array of numbers representing pixels of the image, with each number in the array including a value between 0 and 255 representing the pixel intensity at that location in the array. Using the previous example from above, the array can include a 28×28×3 array of numbers, with 28 rows and 28 columns of pixels and three color components (e.g., a red component, a green component, and a blue component, or a luma component and two chroma components, etc.). The image is passed through a convolutional hidden layer 922 a, an optional nonlinear activation layer, a pooling hidden layer 922 b, and a fully connected hidden layer 922 c to obtain an output at the output layer 924. While only one of each hidden layer is shown in FIG. 9, those skilled in the art will appreciate that multiple convolutional hidden layers, nonlinear layers, pooling hidden layers, and / or fully connected layers can be included in the CNN 900. As previously explained, the output can indicate a single class of object, or can include probabilities of the class that best describes the object in the image.
[0121] The first layer of the CNN 900 is a convolutional hidden layer 922a. The convolutional hidden layer 922a analyzes the image data in the input layer 920. Each node in the convolutional hidden layer 922a is connected to a region of nodes (pixels) in the input image, called the receptive field. The convolutional hidden layer 922a can be thought of as one or more filters (each filter corresponds to a different activation or feature map), and each convolutional iteration of a filter is a node or neuron in the convolutional hidden layer 922a. For example, the region of the input image that a filter is responsible for in each convolutional iteration is the receptive field of the filter. In one illustrative example, if the input image contains a 28x28 array and each filter (and corresponding receptive field) is a 5x5 array, there will be 24x24 nodes in the convolutional hidden layer 922a. As each node learns to analyze a specific local receptive field in the input image, each connection between a node and the receptive field for that node learns a weight, and possibly an overall bias. Each node in the hidden layer 922a has the same weights and biases (called shared weights and shared biases). For example, a filter has an array of weights (numbers) and a depth equal to the input. The filter has a depth of 3 in the example video frame (according to the three color components of the input image). An illustrative example size of the filter array is 5x5x3, which corresponds to the size of the receptive field of the node.
[0122] The convolutional nature of the convolutional hidden layer 922a results from the fact that each node in the convolutional layer is applied to its corresponding receptive field. For example, the filter in the convolutional hidden layer 922a can start in the upper left corner of the input image array and convolve around the input image. As described above, each convolutional iteration of the filter can be considered a node or neuron in the convolutional hidden layer 922a. In each convolutional iteration, the value of the filter is multiplied by a corresponding number of original pixel values of the image (e.g., a 5x5 filter array is multiplied by a 5x5 array of input pixel values in the upper left corner of the input image array). The multiplications from each convolutional iteration can be added together to obtain a total for that iteration or node. The process then continues at the next location in the input image according to the receptive field of the next node in the convolutional hidden layer 922a. For example, the filter can be moved a certain step amount to the next receptive field. This step amount can be set to 1 or another appropriate amount. For example, if the step amount is set to 1, the filter is moved one pixel to the right in each convolution iteration. Processing the filter at each unique location in the input volume produces a number representing the filter result for that location, from which a total value is determined for each node in the convolutional hidden layer 922a.
[0123] The mapping from the input layer to the convolutional hidden layer 922a is called an activation map (or feature map). The activation map contains a value for each node that represents the filter result at each location in the input volume. The activation map may include an array containing various aggregate values resulting from each iteration of the filter on the input volume. For example, if a 5x5 filter is applied to each pixel (in steps of 1) of a 28x28 input image, the activation map would include a 24x24 array. The convolutional hidden layer 922a can include several activation maps to identify multiple features in an image. The example shown in Figure 9 includes three activation maps. Using the three activation maps, the convolutional hidden layer 922a can detect three different types of features, each detectable across the entire image.
[0124] In some examples, a nonlinear hidden layer may be applied after the convolutional hidden layer 922a. A nonlinear layer may be used to introduce nonlinearity into a system that previously computed a linear operation. One illustrative example of a nonlinear layer is a rectified linear unit (ReLU) layer. A ReLU layer may apply a function f(x)=max(0,x) to all of the values in the input volume, which changes all negative activations to 0. Thus, ReLU can enhance the nonlinear nature of the network 900 without affecting the receptive field of the convolutional hidden layer 922a.
[0125] A pooling hidden layer 922b may be applied after the convolutional hidden layer 922a (and after the nonlinear hidden layer, if used). The pooling hidden layer 922b is used to simplify the information in the output from the convolutional hidden layer 922a. For example, the pooling hidden layer 922b may take each activation map output from the convolutional hidden layer 922a and use a pooling function to generate a condensed activation map (or feature map). Max pooling is an example of a function performed by the pooling hidden layer. Other forms of pooling functions, such as average pooling, L2-norm pooling, or other suitable pooling functions, may be used by the pooling hidden layer 922a. A pooling function (e.g., a max pooling filter, an L2-norm filter, or other suitable pooling filter) is applied to each activation map included in the convolutional hidden layer 922a. In the example shown in FIG. 9, three pooling filters are used for the three activation maps in the convolutional hidden layer 922a.
[0126] In some examples, maximization pooling may be used by applying a maximization pooling filter (e.g., having a size of 2×2) with a step amount (e.g., equal to the dimension of the filter, such as a step amount of 2) to the activation map output from the convolutional hidden layer 922a. The output from the maximization pooling filter contains the largest number of any subregion around which the filter convolves. Using a 2×2 filter as an example, each unit in the pooling layer can summarize a region of 2×2 nodes (each node is a value in the activation map) in the previous layer. For example, four values (nodes) in the activation map are analyzed by 2×2 maximization pooling at each iteration of the filter, and the maximum of the four values is output as the “max” value. If such a maximization pooling filter is applied to the activation filter from the convolutional hidden layer 922a with a dimension of 24×24 nodes, the output from the pooling hidden layer 922b is an array of 12×12 nodes.
[0127] In some examples, an L2 norm pooling filter may also be used, which involves calculating the square root of the sum of the squares of the values in a 2×2 region (or other suitable region) of the activation map (rather than calculating the maximum value as is done in max pooling) and using the calculated value as the output.
[0128] Intuitively, a pooling function (e.g., max pooling, L2 norm pooling, or other pooling function) determines whether a given feature is found anywhere within a region of the image and discards the exact location information. This can be done without affecting the outcome of feature detection because, once the feature has been found, the exact location of the feature is not as important as its approximate location relative to other features. Max pooling (as well as other pooling methods) offers the advantage of having far fewer features pooled, thus reducing the number of parameters required in later layers of CNN 900.
[0129] The final layer of connections in the network is a fully connected layer that connects every node from the pooling hidden layer 922b to every output node in the output layer 924. Using the example above, the input layer includes 28x28 nodes that encode pixel intensities of the input image, the convolutional hidden layer 922a includes 3x24x24 hidden feature nodes based on applying 5x5 local receptive fields (for the filters) to three activation maps, and the pooling layer 922b includes a layer of 3x12x12 hidden feature nodes based on applying a maximization pooling filter to a 2x2 region across each of the three feature maps. Extending this example, the output layer 924 could include 10 output nodes. In such an example, every node in the 3x12x12 pooling hidden layer 922b is connected to every node in the output layer 924.
[0130] The fully connected layer 922c can take the output of the previous pooling layer 922b (which should represent the activation map of high-level features) and determine the features that are most correlated to a particular class. For example, the fully connected layer 922c can determine the high-level features that are most strongly correlated to a particular class and can include weights (nodes) for the high-level features. The weights of the fully connected layer 922c can be multiplied by the weights of the pooling hidden layer 922b to obtain probabilities for different classes. For example, if the CNN 900 is being used to predict that an object in a video frame is a person, there will be high values in the activation map that represent high-level features of a person (e.g., having two legs, having a face at the top of the object, having two eyes at the top left and top right of the face, a nose in the center of the face, a mouth at the bottom of the face, and / or other features common to people).
[0131] In some examples, the output from the output layer 924 can include an M-dimensional vector (in the previous example, M=10), where M can include the number of classes the program must choose from when classifying objects in an image. Other exemplary outputs can also be provided. Each number in the M-dimensional vector can represent the probability that the object is of a certain class. In one illustrative example, if a 10-dimensional output vector represents 10 different classes of objects as [0 0 0.05 0.8 0 0.15 0 0 0 0], the vector indicates that there is a 5% probability that the image is of the third class of object (e.g., a dog), an 80% probability that the image is of the fourth class of object (e.g., a person), and a 15% probability that the image is of the sixth class of object (e.g., a kangaroo). The probability of a class can be considered a confidence level that the object is part of that class.
[0132] 10 is a diagram illustrating an example of a system for implementing some aspects of the present technology. Specifically, FIG. 10 illustrates an example of a computing system 1000, which may be, for example, an internal computing system, a remote computing system, a camera, or any computing device comprising any of these components, the components of the system communicating with each other using a connection 1005. The connection 1005 may be a physical connection using a bus or a direct connection to a processor 1010, such as in a chipset architecture. The connection 1005 may also be a virtual connection, a network connection, or a logical connection.
[0133] In some examples, computing system 1000 is a distributed system in which the functionality described in this disclosure may be distributed across a data center, multiple data centers, a peer network, etc. In some examples, one or more of the described system components represent multiple components, each performing some or all of the functionality covered by the component description. In some cases, the components may be physical or virtual devices.
[0134] The exemplary system 1000 includes at least one processing unit (CPU or processor) 1010 and connections 1005 coupling various system components to the processor 1010, including system memory 1015 such as read-only memory (ROM) 1020 and random access memory (RAM) 1025. The computing system 1000 may include a cache 1012 of high-speed memory directly connected to the processor 1010, in close proximity to the processor 1010, or integrated as part of the processor 1010.
[0135] Processor 1010 may include any general-purpose processor and hardware or software services, such as services 1032, 1034, and 1036 stored on storage device 1030, configured to control processor 1010 as well as special-purpose processors where software instructions are incorporated into the actual processor design. Processor 1010 may essentially be a completely self-contained computing system, including multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors may be symmetric or asymmetric.
[0136] To enable user interaction, computing system 1000 includes input devices 1045, which may represent any number of input mechanisms, such as a microphone for speech, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, speech, etc. Computing system 1000 may also include output devices 1035, which may be one or more of a number of output mechanisms. In some instances, a multimodal system may allow a user to provide multiple types of input / output to communicate with computing system 1000. Computing system 1000 may include a communication interface 1040, which may generally govern and manage user input and system output.Communication interfaces include audio jacks / plugs, microphone jacks / plugs, Universal Serial Bus (USB) ports / plugs, Apple® Lightning® ports / plugs, Ethernet ports / plugs, fiber optic ports / plugs, proprietary wired ports / plugs, BLUETOOTH® wireless signal transmission, BLUETOOTH® low energy (BLE) wireless signal transmission, IBEACON® wireless signal transmission, Radio Frequency Identification (RFID) wireless signal transmission, Near Field Communication (NFC) wireless signal transmission, Dedicated Short Range Communication (DSRC) wireless signal transmission, 802.10 Wi-Fi wireless signal transmission, Wireless Local Area Network (WLAN) signal transmission, Visible Light Communication (VLC), and Worldwide Interoperability for Microwave Access The communication interface 1040 may perform or facilitate the reception and / or transmission of wired or wireless communications using wired and / or wireless transceivers, including those utilizing WiMAX (Wireless Maximum Access Point), infrared (IR) communications wireless signal transmissions, public switched telephone network (PSTN) signal transmissions, integrated services digital network (ISDN) signal transmissions, 3G / 4G / 5G / LTE cellular data network wireless signal transmissions, ad hoc network signal transmissions, radio wave signal transmissions, microwave signal transmissions, infrared signal transmissions, visible light signal transmissions, ultraviolet light signal transmissions, wireless signal transmissions along the electromagnetic spectrum, or any combination thereof. The communication interface 1040 may also include one or more Global Navigation Satellite System (GNSS) receivers or transceivers used to determine the location of the computing system 1000 based on reception of one or more signals from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the United States' Global Positioning System (GPS), the Russian Global Navigation Satellite System (GLONASS), the Chinese BeiDou Navigation Satellite System (BDS), and the European Galileo GNSS.Since there is no restriction to operating on any particular hardware configuration, the basic features herein may be easily replaced by improved hardware or firmware configurations as they are developed.
[0137] The storage device 1030 may be a non-volatile and / or non-transitory and / or computer readable memory device, such as a magnetic cassette, a flash memory card, a solid state memory device, a digital versatile disk, a cartridge, a floppy disk, a flexible disk, a hard disk, a magnetic tape, a magnetic strip / stripe, any other magnetic storage medium, a flash memory, a memristor memory, any other solid state memory, a compact disc read only memory (CD-ROM) optical disk, a rewritable compact disc (CD) optical disk, a digital video disc (DVD) optical disk, a Blu-ray Disc (BDD) optical disk, a holographic optical disk, another optical medium, a secure digital (SD) card, a micro secure digital (microSD) card, a memory stick card, a smart card chip, an EMV chip, a subscriber It may be a hard disk or other type of computer-readable medium capable of storing data that is accessible by a computer, such as an identification module (SIM) card, a mini / micro / nano / pico SIM card, another integrated circuit (IC) chip / card, random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash EPROM (FLASHEPROM), cache memory (L1 / L2 / L3 / L4 / L5 / L#), resistive random access memory (RRAM / ReRAM), phase change memory (PCM), spin-transfer torque RAM (STT-RAM), another memory chip or cartridge, and / or a combination thereof.
[0138] The storage device(s) 1030 may include software services, servers, services, etc. that cause the system to perform functions when code defining such software is executed by the processor(s) 1010. In some examples, hardware services that perform particular functions may include software components stored on computer-readable media that interface with the necessary hardware components, such as the processor(s) 1010, connections 1005, output devices 1035, etc., to perform the functions.
[0139] As used herein, the term “computer-readable medium” includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or transporting instructions and / or data. Computer-readable media may also include non-transitory media that can store data and do not include carrier waves and / or transitory electronic signals propagating wirelessly or over wired connections. Examples of non-transitory media may include, but are not limited to, magnetic disks or tapes, optical storage media such as compact discs (CDs) or digital versatile discs (DVDs), flash memory, memory, or memory devices. A computer-readable medium may store code and / or machine-executable instructions, which may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted using any suitable means including memory sharing, message passing, token passing, network transmission, etc.
[0140] In some examples, computer-readable storage devices, media, and memories may include cables or wireless signals containing bitstreams, etc. However, when referred to, non-transitory computer-readable storage media specifically excludes media such as energy, carrier signals, electromagnetic waves, and signals in its own right.
[0141] Specific details are provided in the above description to provide a thorough understanding of the examples provided herein. However, it will be understood by those skilled in the art that the examples may be practiced without these specific details. For clarity of explanation, in some instances, the technology may be presented as including individual functional blocks, including functional blocks that include devices, device components, operations, steps, or routines in methods embodied in software or a combination of hardware and software. Additional components other than those shown in the figures and / or described herein may be used. For example, circuits, systems, networks, processes, or other components may be shown as components in block diagram form to avoid obscuring the examples in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail to avoid obscuring the examples.
[0142] Individual examples may be described above as a process or method that is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. While a flowchart may describe operations as a sequential process, many of the operations may be performed in parallel or simultaneously. In addition, the order of operations may be rearranged. A process terminates when its operations are completed, but may have additional operations not included in the diagram. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination may correspond to the function returning to the calling function or the main function.
[0143] The processes and methods according to the examples described above can be implemented using computer-executable instructions stored or otherwise available from a computer-readable medium. Such instructions can include, for example, instructions and data that cause a general-purpose computer, special-purpose computer, or processing device to perform, or otherwise configure a general-purpose computer, special-purpose computer, or processing device to perform, a particular function or group of functions. Portions of the computer resources used can be accessible over a network. The computer-executable instructions can be, for example, binary, intermediate-format instructions such as assembly language, firmware, source code, etc. Examples of computer-readable media that may be used to store instructions, information used, and / or information created during methods according to the described examples include magnetic or optical disks, flash memory, USB devices with non-volatile memory, network-attached storage devices, etc.
[0144] Devices implementing processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, the program code or code segments (e.g., a computer program product) to perform the necessary tasks may be stored on a computer-readable or machine-readable medium. A processor may perform the necessary tasks. Typical examples of form factors include laptops, smartphones, mobile phones, tablet devices or other small form factor personal computers, personal digital assistants, rack-mounted devices, standalone devices, etc. The functionality described herein may also be embodied in a peripheral device or an add-in card. Such functionality may also be implemented on a circuit board, in different chips, or different processes running in a single device, as further examples.
[0145] The instructions, media for carrying such instructions, computing resources for executing the instructions, and other structures for supporting such computing resources are exemplary means for providing the functionality described in this disclosure.
[0146] In the foregoing description, aspects of the present application have been described with reference to specific examples thereof, but those skilled in the art will recognize that the present application is not limited thereto. Accordingly, while illustrative examples of the present application have been described in detail herein, it should be understood that the concepts of the present invention may be embodied and employed in various other ways, and that the appended claims are intended to be construed to include such variations except insofar as limited by the prior art. Various features and aspects of the application examples described above may be used individually or together. Moreover, the examples may be used in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of the present specification. Accordingly, the specification and drawings should be regarded as illustrative and not restrictive. For purposes of illustration, methods have been described in a particular order. It should be appreciated that, in alternative examples, methods may be performed in an order different from that described.
[0147] Those skilled in the art will appreciate that the less than ("<") and greater than (">") symbols or terms used herein may be replaced with the less than or equal to ("≦") and greater than or equal to ("≧") symbols, respectively, without departing from the scope of this description.
[0148] When a component is described as being "configured to" perform some operations, such configuration can be achieved, for example, by designing electronic circuitry or other hardware to perform the operations, by programming programmable electronic circuitry (e.g., a microprocessor or other suitable electronic circuitry) to perform the operations, or any combination thereof.
[0149] The phrase "coupled to" refers to any component that is physically connected to another component, either directly or indirectly, and / or that is in communication with another component, either directly or indirectly (e.g., connected to another component via a wired or wireless connection and / or other suitable communication interface).
[0150] Claim language or other language reciting "at least one of" a set and / or "one or more" of a set indicates that one element of the set or multiple elements of the set (in any combination) satisfies the claim. For example, claim language reciting "at least one of A and B" means A, B, or A and B. As another example, claim language reciting "at least one of A, B, and C" means A, B, C, or A and B, or A and C, or B and C, or A, B, and C. The language "at least one of" a set and / or "one or more" of a set does not limit the set to the items listed in the set. For example, claim language reciting "at least one of A and B" can mean A, B, or A and B, and can additionally include items not listed in the set of A and B.
[0151] The various illustrative logical blocks, modules, circuits, and algorithmic operations described in connection with the examples disclosed herein may be implemented as electronic hardware, computer software, firmware, or combinations thereof. To clearly illustrate this interchangeability of hardware and software, the various illustrative components, blocks, modules, circuits, and operations have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.
[0152] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, such as a general-purpose computer, a wireless communication device handset, or an integrated circuit device having multiple uses, including applications in wireless communication device handsets and other devices. Any features described as modules or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be realized at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may comprise a memory or data storage medium such as a random access memory (RAM), e.g., a synchronous dynamic random access memory (SDRAM), a read-only memory (ROM), a nonvolatile random access memory (NVRAM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a magnetic or optical data storage medium, or the like. The techniques may additionally or alternatively be realized at least in part by a computer-readable communications medium, such as a propagated signal or wave, that carries or communicates program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer.
[0153] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Such a processor may be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor, but alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Accordingly, the term “processor,” as used herein, may refer to any of the above structures, any combination of the above structures, or any other structure or apparatus suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided in dedicated software or hardware modules configured for encoding and decoding, or incorporated into a composite video encoder / decoder (codec).
[0154] Exemplary aspects of the present disclosure are as follows.
[0155] Aspect 1: An apparatus for capturing self-images within an extended reality environment, the apparatus including: a memory; and one or more processors coupled to the memory, the one or more processors configured to: capture a pose of a user of an extended reality system, the user's pose including a location of the user within a real-world environment associated with the extended reality system; generate a digital representation of the user, the digital representation of the user reflecting the user's pose; capture one or more frames of the real-world environment; and overlay the digital representation of the user on the one or more frames of the real-world environment.
[0156] Aspect 2: The device of claim 1, wherein the one or more processors are configured to overlay a digital representation of the user onto one or more frames of the real-world environment in a frame location corresponding to the location of the user within the real-world environment.
[0157] Aspect 3: The device of claim 1, wherein the one or more processors are configured to generate a digital representation of the user prior to capturing one or more frames of the real-world environment.
[0158] Aspect 4: The device of any one of claims 1 to 3, wherein the one or more processors are configured to display within a display of the extended reality system, through which the real-world environment is viewed, a digital representation of the user within a display location corresponding to the user's location within the real-world environment.
[0159] Aspect 5: The device of claim 4, wherein the one or more processors are configured to detect user input corresponding to instructions to capture one or more frames of the real-world environment while a digital representation of the user is displayed within a display of the extended reality system, and capture one or more frames of the real-world environment based on the user input.
[0160] Aspect 6: The device of claim 1, wherein the one or more processors are configured to capture one or more frames of the real-world environment before capturing the user's pose.
[0161] Aspect 7: The device of claim 6, wherein the one or more processors are configured to display within a display of the extended reality system a digital representation of the user within a display location corresponding to the user's location within the real-world environment, the digital representation corresponding to the user's location within the real-world environment.
[0162] Aspect 8: The device of claim 7, wherein the one or more processors are configured to update a display location of the digital representation of the user based on detecting a change in the user's location within the real-world environment.
[0163] Aspect 9: The device of claim 7, wherein the one or more processors are further configured to: detect a user input corresponding to an instruction to capture a pose of the user while the digital representation of the user is displayed within the display of the extended reality system; and capture the pose of the user based on the user input.
[0164] Aspect 10: The apparatus of any one of claims 1 to 9, wherein the one or more processors are configured to generate a first digital representation of the user of a first fidelity and obtain a second digital representation of the user of a second fidelity, the second fidelity being higher than the first fidelity.
[0165] Aspect 11: The device of claim 10, wherein the one or more processors are configured to: display a first digital representation of the user within a display of the extended reality system before a pose of the user is captured; generate a second digital representation of the user based on the pose of the user being captured; and overlay the second digital representation of the user onto one or more frames of the real-world environment.
[0166] Aspect 12: The device of claim 10, wherein the one or more processors are configured to: display a first digital representation of the user within a display of the extended reality system before one or more frames of the real-world environment are captured; generate a second digital representation of the user based on the one or more frames of the real-world environment being captured; and overlay the second digital representation of the user on the one or more frames of the real-world environment.
[0167] Aspect 13: The apparatus of claim 10, wherein the first digital representation is based on a first machine learning algorithm and the second digital representation of the user is based on a second machine learning algorithm.
[0168] Aspect 14: The apparatus of claim 13, wherein the one or more processors are configured to generate a first digital representation of the user based on implementing a first machine learning algorithm on the extended reality system, and to cause a server configured to generate the digital representation of the user to generate a second digital representation of the user based on implementing a second machine learning algorithm.
[0169] Aspect 15: The device of any one of claims 1 to 14, wherein the one or more processors are configured to capture a pose of a person in a real-world environment; generate a digital representation of the person, where the digital representation of the person reflects the pose of the person; and overlay the digital representation of the user and the digital representation of the person onto one or more frames of the real-world environment.
[0170] Aspect 16: The device of claim 15, wherein the one or more processors are configured to generate a digital representation of the person based at least in part on information associated with the digital representation of the person received from the extended reality system.
[0171] Aspect 17: The device of claim 16, wherein the information associated with the digital representation of the person includes a machine learning model trained to generate the digital representation of the person.
[0172] Aspect 18: The device of any one of claims 1 to 17, wherein the one or more processors are configured to capture a plurality of poses of the user associated with a plurality of frames, generate a plurality of digital representations of the user corresponding to the plurality of frames, and overlay the plurality of digital representations of the user onto one or more frames of a real-world environment, wherein the one or more frames of the real-world environment comprise a plurality of frames of the real-world environment.
[0173] Aspect 19: The device of any one of claims 1 to 18, wherein the one or more processors are configured to generate a digital representation of the user using a first machine learning algorithm and to overlay the digital representation of the user onto one or more frames of a real-world environment using a second machine learning algorithm.
[0174] Aspect 20: The device of any one of claims 1 to 19, wherein the one or more processors are configured to capture a pose of the user based at least in part on image data captured by an inward-facing camera system of the extended reality system.
[0175] Aspect 21: The device of any one of claims 1 to 20, wherein the one or more processors are configured to capture a pose of the user based at least in part on determining a facial expression of the user.
[0176] Aspect 22: The device of any one of claims 1 to 21, wherein the one or more processors are configured to capture a pose of the user based at least in part on determining a gesture of the user.
[0177] Aspect 23: The device of any one of claims 1 to 22, wherein the one or more processors are configured to determine a location of the user within the real-world environment based at least in part on generating a three-dimensional map of the real-world environment.
[0178] Aspect 24: The device of any one of claims 1 to 23, wherein the one or more processors are configured to capture one or more frames of the real-world environment using an outward-facing camera system of the extended reality system.
[0179] Aspect 25: The device of any one of claims 1 to 24, wherein the device comprises an extended reality system.
[0180] Aspect 26: The apparatus of any one of claims 1 to 25, wherein the apparatus comprises a mobile device.
[0181] Aspect 27: The device of any one of claims 1 to 26, further comprising a display.
[0182] Aspect 28: A method for capturing a self-image within an extended reality environment, the method comprising the steps of capturing a pose of a user of an extended reality system, the user's pose including the user's location within a real-world environment associated with the extended reality system; generating a digital representation of the user, the digital representation of the user reflecting the user's pose; capturing one or more frames of the real-world environment; and overlaying the digital representation of the user onto the one or more frames of the real-world environment.
[0183] Aspect 29: The method of claim 28, wherein overlaying the digital representation of the user onto one or more frames of the real-world environment includes overlaying the digital representation of the user within a frame location corresponding to the user's location within the real-world environment.
[0184] Aspect 30: The method of claim 28, wherein the step of generating a digital representation of the user is performed before capturing one or more frames of the real-world environment.
[0185] Aspect 31: The method of any one of claims 28 to 30, further comprising displaying within a display of the extended reality system, a digital representation of the user within a display location corresponding to the user's location within the real-world environment, through which the real-world environment is viewed.
[0186] Aspect 32: The method of claim 31, wherein the step of capturing one or more frames of the real-world environment further includes the steps of: detecting user input corresponding to an instruction to capture one or more frames of the real-world environment while the digital representation of the user is displayed within the display of the extended reality system; and capturing one or more frames of the real-world environment based on the user input.
[0187] Aspect 33: The method of claim 28, wherein the step of capturing one or more frames of the real-world environment is performed before capturing the user's pose.
[0188] Aspect 34: The method of claim 33, further comprising displaying a digital representation of the user in a display location corresponding to the user's location in the real-world environment within a display of the extended reality system on which one or more frames of the real-world environment are displayed.
[0189] Aspect 35: The method of claim 34, further comprising updating a display location of the digital representation of the user based on detecting a change in the user's location within the real-world environment.
[0190] Aspect 36: The method of claim 34, wherein the step of capturing a pose of a user of the extended reality system further includes the steps of detecting a user input corresponding to an instruction to capture a pose of the user while a digital representation of the user is displayed within a display of the extended reality system, and capturing the pose of the user based on the user input.
[0191] Aspect 37: The method of any one of claims 28 to 36, wherein generating a digital representation of the user includes generating a first digital representation of the user of a first fidelity and obtaining a second digital representation of the user of a second fidelity, the second fidelity being higher than the first fidelity.
[0192] Aspect 38: The method of claim 37, further comprising the steps of: displaying a first digital representation of the user within a display of the extended reality system before the user's pose is captured; generating a second digital representation of the user based on the user's pose being captured; and overlaying the second digital representation of the user onto one or more frames of the real-world environment.
[0193] Aspect 39: The method of claim 37, further comprising the steps of: displaying a first digital representation of the user within a display of the extended reality system before one or more frames of the real-world environment are captured; generating a second digital representation of the user based on the one or more frames of the real-world environment being captured; and overlaying the second digital representation of the user on the one or more frames of the real-world environment.
[0194] Aspect 40: The method of claim 37, wherein the first digital representation is based on a first machine learning algorithm and the second digital representation of the user is based on a second machine learning algorithm.
[0195] Aspect 41: The method of claim 40, wherein generating a first digital representation of the user includes implementing a first machine learning algorithm on the extended reality system, and obtaining a second digital representation of the user includes causing a server configured to generate digital representations of users to generate the second digital representation of the user based on implementing the second machine learning algorithm.
[0196] Aspect 42: The method of any one of claims 28 to 41, further comprising the steps of capturing a pose of a person in a real-world environment, generating a digital representation of the person, wherein the digital representation of the person reflects the pose of the person, and overlaying the digital representation of the user and the digital representation of the person onto one or more frames of the real-world environment.
[0197] Aspect 43: The method of claim 42, wherein the digital representation of the person is generated based at least in part on information associated with the digital representation of the person received from the person's extended reality system.
[0198] Aspect 44: The method of claim 43, wherein the information associated with the digital representation of the person includes a machine learning model trained to generate the digital representation of the person.
[0199] Aspect 45: The method of any one of claims 28 to 44, further comprising the steps of capturing a plurality of poses of the user associated with a plurality of frames, generating a plurality of digital representations of the user corresponding to the plurality of frames, and overlaying the plurality of digital representations of the user onto one or more frames of a real-world environment, wherein the one or more frames of the real-world environment comprise a plurality of frames of the real-world environment.
[0200] Aspect 46: The method of any one of claims 28 to 45, wherein generating a digital representation of the user includes using a first machine learning algorithm, and overlaying the digital representation of the user onto one or more frames of the real-world environment includes using a second machine learning algorithm.
[0201] Aspect 47: The method of any one of claims 28 to 46, wherein capturing the user's pose includes capturing the image data using an inward-facing camera system of the extended reality system.
[0202] Aspect 48: The method of any one of claims 28 to 47, wherein capturing a pose of the user includes determining a facial expression of the user.
[0203] Aspect 49: The method of any one of claims 28 to 48, wherein capturing a pose of the user includes determining a gesture of the user.
[0204] Aspect 50: The method of any one of claims 28 to 49, further comprising determining a location of the user within the real-world environment based at least in part on generating a three-dimensional map of the real-world environment.
[0205] Aspect 51: The method of any one of claims 28 to 50, wherein capturing one or more frames of the real-world environment includes capturing the image data using an outward-facing camera system of the extended reality system.
[0206] Aspect 52: A non-transitory computer-readable storage medium for capturing self-images within an extended reality environment, the non-transitory computer-readable storage medium including instructions stored therein that, when executed by one or more processors, cause the one or more processors to perform operations according to any of aspects 1 to 51.
[0207] Aspect 53: An apparatus for capturing a self-image within an extended reality environment, the apparatus comprising means for performing the operations according to any of aspects 1 to 51. [Explanation of symbols]
[0208] 102 Selfie Picture Frames 104 Selfie Picture Frames 106 Background Frames 108 Avatar 110 Positioning Tools 116 frames 118 Avatar 120 Avatar 200 Extended Reality System 202 Image Sensor 204 Accelerometer 206 Gyroscope 208 Storage 210 Computational Components 212 Central Processing Unit (CPU) 214 Graphics Processing Unit (GPU) 216 Digital Signal Processor (DSP) 218 Image Signal Processor (ISP) 220 Extended Reality (XR) Engine 222 Selfie Image Engine 224 Image Processing Engine 226 Rendering Engine 300(A) Selfie Image Capture System 300(B) Selfie Image Capture System 302 Self Image Start Engine 304 Avatar Engine 304(A) Avatar Engine 304(B) Avatar Engine 306 Background Frame Engine 308 Configuration Engine 310 User Input 312 User Pose 314 Background Frame 316 Selfie Picture Frames 318 Avatar 318(A) Avatar 318(B) Avatar 402 frames 404 Frame 406 frames 408 frames 410 frames 412 frames 414 Preview Window 500(A) Process 500(B) Process 600(A) Process 600(B) Process 622 Multi-User Selfie Image Capture System 624(1)~624(N) Avatar Network 626 Selfie Image Generator 628 multi-user selfies 630(1)~(N) User pause 700 processes 800 Deep Learning Neural Networks 820 Input Layer 822a Hidden layer 822b Hidden layer 822n hidden layer 824 output layer 826 nodes 900 Convolutional Neural Networks 920 Input Layer 922a Convolutional Hidden Layer 922b Pooling hidden layer 922c Fully connected hidden layer 924 output layer 1000 Computing Systems 1005 connections 1010 processor 1012 Cache 1015 system memory 1020 Read-Only Memory (ROM) 1025 Random Access Memory (RAM) 1030 Storage Devices 1032 Service 1034 Service 1035 output device 1036 Service 1040 Communication Interface 1045 Input Devices
Claims
1. 1. An apparatus for capturing a self-image within an extended reality environment, comprising: Memory and one or more processors coupled to the memory, capturing a pose of a user of an extended reality system, the pose of the user including a location of the user within a portion of a real-world environment associated with the extended reality system; generating a digital avatar representation of the user, wherein the digital avatar representation of the user reflects the pose of the user; and capturing one or more frames of the portion of the real-world environment without the user being present in the one or more frames; superimposing the digital avatar representation of the user onto the one or more frames of the portion of the real-world environment in which the user is not present in the one or more frames in frame locations corresponding to the location of the user in the portion of the real-world environment associated with the captured pose, the digital avatar representation being stationary as the user moves through the real-world environment to capture the one or more frames of the portion of the real-world environment in which the user is not present in the one or more frames; configured to: wherein the one or more processors are configured to capture the one or more frames of the portion of the real-world environment before capturing the pose of the user, the pose of the user further including limb positions or hand gestures of the user based on image data captured by a camera used to capture the one or more frames of the portion of the real-world environment, and the digital avatar representation of the user is overlaid on the one or more frames at a location corresponding to the user's location in the real-world environment at the time the pose of the user was captured.
2. 2. The device of claim 1, wherein, to overlay the digital avatar representation of the user onto the one or more frames, the one or more processors are configured to display the digital avatar representation of the user in the frame location corresponding to the location of the user in the portion of the real-world environment within a display of the extended reality system on which the one or more frames of the portion of the real-world environment are displayed.
3. 3. The device of claim 2, wherein the one or more processors are configured to update the frame location of the digital avatar representation of the user based on detecting a change in the location of the user within the portion of the real-world environment.
4. the one or more processors: detecting a user input corresponding to a command to capture the pose of the user while the digital avatar representation of the user is displayed within the display of the extended reality system; capturing the pose of the user based on the user input; The apparatus of claim 2 , further configured to:
5. 1. A method for capturing a self-image in an extended reality environment, comprising: capturing a pose of a user of an extended reality system, the pose of the user including a location of the user within a portion of a real-world environment associated with the extended reality system; generating a digital avatar representation of the user, wherein the digital avatar representation of the user reflects the pose of the user; capturing one or more frames of the portion of the real-world environment without the user being present in the one or more frames; superimposing the digital avatar representation of the user onto the one or more frames of the portion of the real-world environment in which the user is not present in the one or more frames in a frame location corresponding to the location of the user in the portion of the real-world environment associated with the captured pose, the digital avatar representation remaining stationary as the user moves through the real-world environment to capture the one or more frames of the portion of the real-world environment in which the user is not present in the one or more frames; Including, wherein the step of capturing the one or more frames of the portion of the real-world environment is performed before capturing the pose of the user, the pose of the user further including limb positions or hand gestures of the user based on image data captured by a camera used to capture the one or more frames of the portion of the real-world environment, and the digital avatar representation of the user is overlaid on the one or more frames at a location corresponding to the user's location in the real-world environment at the time the pose of the user was captured.
6. 6. The method of claim 5, wherein overlaying the digital avatar representation of the user onto the one or more frames includes displaying the digital avatar representation of the user in the frame location corresponding to the location of the user in the portion of the real-world environment within a display of the extended reality system on which the one or more frames of the portion of the real-world environment are displayed.
7. 7. The method of claim 6, further comprising updating the frame location of the digital avatar representation of the user based on detecting a change in the location of the user within the portion of the real-world environment.
8. Capturing the pose of the user of the extended reality system includes: detecting a user input corresponding to a command to capture the pose of the user while the digital avatar representation of the user is displayed within the display of the extended reality system; capturing the pose of the user based on the user input; The method of claim 6 further comprising:
Citation Information
Patent Citations
Human body gesture-based region and volume selection for hmd
JP2016514298A
Image display system, head-mounted display control device, and operating method and operating program for same
WO2018012207A1
Video synthesis device, video synthesis method and recording medium
WO2020110323A1
Detecting input in artificial reality systems based on a pinch and pull gesture
WO2020247279A1