Systems and methods for predicting lower body posture
By using a generative adversarial network (GAN) model, combined with sensor data and image data, the problem of accuracy in generating lower body poses in virtual reality systems was solved, achieving realistic replication of the user's lower body poses and improving the generation effect of full-body poses.
Patent Information
- Application Number
- CN202180063845.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-09-17
- Filing Date
- 2021-08-16
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2041-08-16
AI Technical Summary
In existing technologies, sensor data and image data have limited accuracy in generating user lower body postures, especially in virtual reality systems, where it is difficult to accurately replicate the user's lower body posture.
By utilizing machine learning models, particularly generative adversarial networks (GANs), corresponding lower body poses are generated from upper body poses. The model is trained by combining sensor data and image data to improve the accuracy of lower body pose generation.
It achieves accurate generation of the user's lower body posture in the virtual reality system, improving the realism and consistency of the whole body posture.
Smart Images

Figure CN116235226B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates generally to predicting a user's lower body pose. BACKGROUND
[0002] Machine learning can be used to enable a machine to automatically detect and process objects. Generally, machine learning typically involves processing a training data set according to a machine learning model and updating the machine learning model based on a training algorithm such that the machine learning model progressively "learns" features in the training data set that are predictive of a desired output. One example of a machine learning model is a neural network, which is a network of interconnected nodes. Groups of nodes can be arranged in layers. A first layer in the network that receives input data can be referred to as an input layer, and a last layer that outputs data from the network can be referred to as an output layer. There can be any number of internal, hidden layers that map nodes in the input layer to nodes in the output layer. In a feedforward neural network, the output of nodes in each layer, other than the output layer, is configured to feed forward into nodes in subsequent layers.
[0003] Artificial reality is a form of reality that has been adjusted in some manner before presentation to a user, which can include, e.g., a virtual reality (VR), an augmented reality (AR), a mixed reality (MR), a hybrid reality, or some combination and / or derivatives thereof. Artificial reality content can include completely generated content or generated content combined with captured content (e.g., a real-world photograph). Artificial reality content can include video, audio, haptic feedback, or some combination thereof, any of which can be presented in a single channel or in multiple channels (such as stereo video that produces a three-dimensional effect to the viewer). Artificial reality can be associated with applications, products, accessories, services, or some combination thereof, that are, e.g., used to create content in an artificial reality and / or used in an artificial reality (such as an activity for performing in an artificial reality). An artificial reality system that provides artificial reality content can be implemented on various platforms, including a head-mounted display (HMD) connected to a host computer system, a standalone HMD, a mobile device or computing system, or any other hardware platform capable of providing artificial reality content to one or more viewers. SUMMARY
[0004] Embodiments described herein relate to methods of generating one or more body poses of a user associated with one or more components of an artificial reality system. To enable a computing system to generate accurate poses of one or more joints that are constituent parts of a body pose, the computing system can receive one or more sensor data or image data from one or more components of the artificial reality system, such as sensor data from motion tracking sensors or image data received from one or more cameras. With various techniques described herein, the image data and sensor data allow the computing system to accurately generate an upper body pose of the user based on the one or more sensor data or image data associated with the artificial reality system. However, while such sensor data and image data are effective for generating an upper body pose of the user, such sensor data and image data are often of limited use in generating a lower body pose of the user.
[0005] To address this issue, particular embodiments described herein utilize a machine learning model to generate a lower body pose in response to receiving a generated upper body pose of the user. In particular embodiments, the machine learning model can be trained to receive a generated upper body pose and generate a corresponding lower body pose. In particular embodiments, the machine learning model can be based on a Generative Adversarial Network (GAN), and the machine learning model can be trained using one or more training poses. Based on these training poses, the machine learning model can learn how to produce realistic lower body poses. Once trained, the machine learning model can receive a generated upper body pose and output a corresponding lower body pose or full body pose that can be used in various applications. For example, the outputted full body pose can be used to generate an avatar of the user in a virtual reality or artificial reality space.
[0006] According to a first aspect of the disclosure, there is provided a method comprising, by a computing system: receiving sensor data captured by one or more sensors coupled to a user; generating, based on the sensor data, an upper body pose corresponding to a first portion of a body of the user, the first portion of the body including a head and arms of the user; generating, by processing the upper body pose using a machine learning model, a lower body pose corresponding to a second portion of the body of the user, the second portion of the body including legs of the user; and generating, based on the upper body pose and the lower body pose, a full body pose of the user.
[0007] The machine learning model can be trained using a second machine learning model trained to determine whether a given full body pose is likely to be generated using the machine learning model.
[0008] The machine learning model can be optimized during training so that a second machine learning model incorrectly determines that a given full-body pose (which is generated using the machine learning model) is unlikely to have been generated using that machine learning model.
[0009] The machine learning model can be trained in the following ways: by processing the second upper body pose using the machine learning model to generate the second lower body pose; by generating the second full-body pose based on the second upper body pose and the second lower body pose; by using the second machine learning model to determine whether the second full-body pose could have been generated using the machine learning model; and by updating the machine learning model based on the determination of the second machine learning model.
[0010] These one or more sensors can be associated with the head-mounted device worn by the user.
[0011] Generating an upper body posture may include: determining multiple postures corresponding to multiple predetermined body parts of the user based on sensor data; and inferring an upper body posture based on the multiple postures.
[0012] These multiple postures may include wrist postures corresponding to the user's wrist. The wrist posture can be determined based on one or more images captured by the head-mounted device worn by the user. The one or more images may depict (1) the user's wrist or (2) the device held by the user.
[0013] The method may further include determining contextual information associated with the time of sensor data acquisition. Lower body posture can be generated by further processing the contextual information using a machine learning model.
[0014] Contextual information can include the application that the user is interacting with when determining the upper body posture.
[0015] The method may also include generating a user avatar based on full-body posture.
[0016] According to a second aspect of this disclosure, one or more computer-readable non-transitory storage media are provided, the one or more computer-readable non-transitory storage media including instructions that, when executed by a server, cause the server to perform the method of the first aspect of this disclosure.
[0017] According to a third aspect of this disclosure, a system is provided, comprising: one or more processors; and one or more computer-readable nontransitory storage media of the second aspect of this disclosure, the one or more computer-readable nontransitory storage media being coupled to the one or more processors.
[0018] The various embodiments disclosed herein are merely examples, and the scope of this disclosure is not limited to these embodiments. Specific embodiments may include, or may not include, all or some of the components, elements, features, functions, operations, or steps of the embodiments disclosed herein. Embodiments according to the invention are specifically disclosed in the appended claims for methods, storage media, and systems, wherein any feature mentioned in one claim class (e.g., method) may also be claimed in another claim class (e.g., system). Dependencies or references in the appended claims are chosen merely for formal reasons. However, protection may also be claimed for any subject matter arising from the intentional reference (in particular multiple dependencies) to any plurality of prior claims, thereby disclosing any combination of the plurality of claims and their features, and any combination of the plurality of claims and their features may be claimed regardless of the dependency chosen in the appended claims. The subject matter for which protection may be claimed includes not only multiple combinations of the plurality of features set forth in the appended plurality of claims, but also any other combination of the plurality of features in the plurality of claims, wherein each feature mentioned in the plurality of claims may be combined with any other feature in the plurality of claims, or with a combination of multiple other features. Furthermore, protection may be claimed in a single claim for any embodiment and feature of the plurality of embodiments and features described or depicted herein, and / or protection may be claimed for any combination of any embodiment and feature of the plurality of embodiments and features described or depicted herein with any embodiment or feature described or depicted herein, or protection may be claimed for any combination of any embodiment and feature of the plurality of embodiments and features described or depicted herein with any feature of the appended claims. Attached Figure Description
[0019] Figure 1 An example artificial reality system is shown.
[0020] Figure 2 Sample body poses associated with the user are shown.
[0021] Figure 3 An example is shown of using a machine learning model to generate a lower body pose by leveraging known predicted upper body poses.
[0022] Figure 4 The configuration for training a machine learning model for lower body posture prediction is shown.
[0023] Figure 5 An example method for training a generator is shown.
[0024] Figure 6This demonstrates how to generate a full-body pose from the generated upper-body pose and the generated lower-body pose.
[0025] Figure 7 An example method for training a discriminator is shown.
[0026] Figure 8 An example method is shown for generating a user's full-body pose based on upper body and lower body poses.
[0027] Figure 9 An example network environment associated with a social networking system is shown.
[0028] Figure 10 An example computer system is shown. Detailed Implementation
[0029] Figure 1 An example artificial reality system 100 is illustrated. In a particular example, artificial reality system 100 may include a head-mounted device 104, a controller 106, and a computing system 108. A user 102 may wear the head-mounted device 104, which may display visual artificial reality content to the user 102. The head-mounted device 104 may include an audio device that may provide audio artificial reality content to the user 102. The head-mounted device 104 may include one or more cameras 110 that may capture images and video of the environment. The head-mounted device 104 may include an eye-tracking system to determine the user 102's vergence distance. Vergence distance may be the distance from the user's eyes to the object where the user's eyes converge (e.g., a real-world object or a virtual object in virtual space). The head-mounted device 104 may be referred to as a head-mounted display (HMD). In a particular example, the computing system 108 may determine the pose of the head-mounted device 104 associated with the user 102. The pose of the head-mounted device can be determined by utilizing any of the sensor data or image data received by the computing system 108.
[0030] One or more controllers 106 may be paired with the artificial reality system 100. In a particular example, the one or more controllers 106 may be equipped with at least one inertial measurement unit (IMU) and an infrared (IR) light-emitting diode (LED) for the artificial reality system 100 to estimate the controller's attitude and / or track its position, enabling a user to perform certain functions via the controller. The one or more controllers 106 may be equipped with one or more trackable markers distributed for tracking by the computing system 108. The one or more controllers 106 may include a touchpad and one or more buttons. The one or more controllers 106 may receive input from a user 102 and forward that input to the computing system 108. The one or more controllers 106 may also provide haptic feedback to the user 102. The computing system 108 may be connected to the head-mounted device 104 and the one or more controllers 106 via a cable or wireless connection. The one or more controllers 106 may include a combination of hardware, software, and / or firmware not explicitly shown herein so as not to obscure other aspects of this disclosure.
[0031] In a particular example, computing system 108 may receive sensor data from one or more sensors of artificial reality system 100. In a particular example, the one or more sensors may be associated with user 102. In a particular example, the one or more sensors may be associated with head-mounted device 104 worn by the user. For example, and not limitingly, head-mounted device 104 may include a gyroscope or inertial measurement unit that tracks the user's real-time movement and outputs sensor data to represent or describe that movement. The sensor data provided by such motion-tracking sensors can be used by VR applications to determine the user's current orientation and provide that orientation to a rendering engine to orient / reorient a virtual camera in 3D space. As another example, and not limitingly, the one or more controllers 106 may include an inertial measurement unit (IMU) and an infrared (IR) light-emitting diode (LED) configured to collect IMU sensor data and send that IMU sensor data to computing system 108. In a particular example, computing system 108 may use one or more tracking techniques (such as, but not limited to, Simultaneous Localization and Mapping (SLAM) tracking or IR-based tracking) to use one or more sensor data to determine the pose of one or more components of artificial reality system 100.
[0032] In a particular example, computing system 108 may receive one or more image data from one or more components of artificial reality system 100. In a particular example, this image data includes image data collected from one or more cameras 110 associated with artificial reality system 100. For example, Figure 1 One or more cameras 110 coupled within a head-mounted device 104 are depicted. The one or more cameras may be positioned to capture one or more images associated with various viewpoints, such as, but not limited to, one or more downward-facing cameras associated with the head-mounted device 104 (e.g., facing the feet of the user 102 when standing).
[0033] In a specific example, computing system 108 can determine the controller pose of one or more controllers 106 associated with user 102. The controller pose associated with user 102 can be determined by utilizing sensor data or image data received by computing system 108 and employing one or more techniques (e.g., but not limited to, computer vision techniques (e.g., image classification)). A method for determining controller pose is further described in U.S. Application No. 16 / 734,172, filed January 3, 2020, entitled "Visual-Inertial Object Tracking with Combined Infrared and Visible Light," the entire contents of which are incorporated herein by reference.
[0034] In a specific example, computing system 108 can control head-mounted device 104 and one or more controllers 106 to provide artificial reality content to user 102 and receive input from user 102. Computing system 108 can be a standalone host computer system, an onboard computer system integrated with head-mounted device 104, a mobile device, or any other hardware platform capable of providing artificial reality content to user 102 and receiving input from user 102.
[0035] Figure 2 A sample full-body pose associated with user 102 is shown. In a particular example, computational system 108 can generate a full-body pose 200 associated with user 102, which includes an upper-body pose 205 and a lower-body pose 215. The full-body pose 200 associated with user 102 can attempt to replicate the position and orientation of one or more joints of user 102 at a specific time using artificial reality system 100. In a particular example, the full-body pose 200 associated with user 102 includes a skeletal framework of inverse kinematics (“skeleton” or “body pose”) that may include a list of one or more joints.
[0036] Upper body posture 205 may correspond to a part of the user 102's body. In a specific example, upper body posture 205 may correspond to a part of the user 102's body, which includes at least the user 102's head and arms. In a specific example, upper body posture 205 may be generated by determining multiple postures corresponding to multiple predetermined body parts or joints of the user 102 (e.g., the user 102's head or wrists). In a specific example, upper body posture 205 may also include, for example, but not limited to, postures of one or more joints associated with the user 102's upper body, such as, but not limited to, head posture 210, wrist posture 220, elbow posture 230, shoulder posture 240, neck posture 250, or upper spine posture 260.
[0037] The lower body posture 215 may correspond to a part of the user 102's body. In a specific example, the lower body posture 215 may correspond to a part of the user 102's body that includes at least the user 102's legs. In a specific example, the lower body posture 215 may also include, for example, but not limited to, joint postures of one or more joints associated with the user 102's lower body, such as, but not limited to, lower spine posture 270, hip posture 280, knee posture 290, or ankle posture 295.
[0038] In a specific example, one or more joint poses can be represented, for example, but not limited to, a subset of parameters representing the position and / or orientation of the individual joints in body pose 200, as a component of full-body pose 200, upper-body pose 205, or lower-body pose 215. The nonlinear solver can parameterize each joint pose associated with user 102 according to the following seven degrees of freedom: three translation values (e.g., x, y, z), three rotation values (e.g., Euler angles in radians), and one uniform scale value. In a specific example, these parameters can be represented using one or more coordinate systems, for example, but not limited to, via an absolute global coordinate system (e.g., x, y, z) or via a local coordinate system relative to a parent joint (e.g., but not limited to, the head joint).
[0039] In a particular example, received sensor data or image data from one or more components of the artificial reality system 100 can be used to determine a full-body pose 200, an upper-body pose 205, or a lower-body pose 215, and one or more joint poses that are components of the full-body pose 200, the upper-body pose 205, or the lower-body pose 215. In a particular example, a combination of one or more techniques can be used to generate one or more joint poses that are components of the full-body pose 200. These techniques include, but are not limited to, localization techniques (e.g., SLAM), machine learning techniques (e.g., neural networks), known spatial relationships with one or more artificial reality system components (e.g., known spatial relationships between head-mounted device 104 and head pose 210), visualization techniques (e.g., image segmentation), or optimization techniques (e.g., nonlinear solvers). In a particular example, one or more of these techniques can be used alone or in combination with one or more other techniques. Such a full-body pose 200 associated with user 102 can be useful for the various applications described herein.
[0040] In a specific example, computing system 108 can generate an upper body posture 205 for user 102. The upper body posture 205 can be generated based at least on sensor data or image data. In a specific example, computing system 108 can determine an upper body posture 200 using one or more postures corresponding to multiple predetermined body parts of user 102, such as, but not limited to, joint postures of one or more joints (e.g., head posture 210 or wrist posture 220) that are components of the upper body posture 205.
[0041] In a particular example, the one or more postures may include a head posture 210 associated with user 102. Head posture 210 may include the position and orientation of the user 102's head joints while wearing the head-mounted device 104. The head posture 210 associated with user 102 may be determined by computing system 108 using any data from sensor data and / or image data received by computing system 108. In a particular example, the head posture 210 associated with user 102 may be determined based on the pose of head-mounted device 104 and the known spatial relationship between head-mounted device 104 and user 102's head. In a particular example, the head posture 210 associated with user 102 may be determined based on sensor data associated with the head-mounted device 104 worn by user 102.
[0042] In a particular example, the one or more postures may include a wrist posture 220 associated with user 102. Wrist posture 220 may include the position and orientation of user 102's wrist joint when interacting with the artificial reality system 100. The wrist posture 220 associated with user 102 may be determined by computing system 108 using any data from sensor data and / or image data received by computing system 108. In a particular example, the image data may include one or more images captured by head-mounted device 104 worn by user 102. The one or more images may depict user 102's wrist or a device held by user 102 (e.g., controller 106). In a particular example, computing system 108 may use one or more computer vision techniques (e.g., but not limited to, image classification or object detection) to determine wrist posture 220 using one or more images. In a particular example, wrist posture 220 may be determined based on controller pose and a known spatial relationship between controller 106 and user 102's wrist.
[0043] In a specific example, the upper body pose 205 can be inferred based on multiple poses (e.g., but not limited to, head pose 210 or wrist pose 220). In a specific example, the upper body pose 205 can be inferred using a nonlinear kinematics optimization solver (“nonlinear solver”), thereby inferring one or more joint poses that are components of the upper body pose 205. In a specific example, the nonlinear solver may include a C++ library built for inverse kinematics, which has a large set of common error functions covering a wide range of applications. In a specific example, the nonlinear solver may provide one or more helper functions for tasks typically associated with global inverse kinematics problems (e.g., joint and skeleton structures, meshes and linear hybrid skins, error functions for common constraints), or it may provide one or more helper functions for mesh deformation (e.g., Laplacian surface deformation), or it may provide one or more I / O functions for various file formats.
[0044] In a specific example, the nonlinear solver can infer one or more joint poses that are part of the upper body pose 205. The one or more joint poses inferred by the nonlinear solver that are part of the upper body pose 205 associated with user 102 can include a kinematic hierarchy at a specific time or state, which is stored as a list of one or more joint poses. In a specific example, the nonlinear solver can include one or more basic solvers supported for inferring one or more joint poses that are part of the upper body pose 205, such as, but not limited to, L-BFGS or Gauss-Newton solvers. In a specific example, the nonlinear solver can utilize a skeleton solver function to solve for the inverse kinematics of a single frame (single body pose). For example, the skeleton solver function can take a current set of one or more parameters (e.g., joint poses) as input and optimize a subset of activated parameters given a defined error function. The convention of nonlinear solvers is to minimize the error of the current function (e.g., the skeleton solver function) in order to infer the accurate upper body pose 205.
[0045] In a specific example, the nonlinear solver can represent one or more joint poses using, for example, but not limited to, a subset of parameters representing the position or orientation of each joint in upper body pose 205. In a specific example, the nonlinear solver can parameterize each joint pose with the following seven degrees of freedom: three translation values (e.g., x, y, z), three rotation values (e.g., Euler angles in radians), and one uniform scaling value. In a specific example, these parameters can be represented using one or more coordinate systems, for example, but not limited to, via an absolute global coordinate system (e.g., x, y, z) or via a local coordinate system relative to a parent joint (e.g., but not limited to, the head joint).
[0046] In a specific example, the nonlinear solver may assign one or more variable limits (e.g., minimum or maximum values) to each joint pose parameter. In a specific example, the nonlinear solver may assign predetermined static weights to each joint pose parameter. These predetermined static weights may be determined, for example, but not limited to, based on the accuracy of the sensor data used to determine the value of each variable. For example, higher static weights may be assigned to joint pose parameters representing head pose 210 and wrist pose 220 associated with user 102 because these joint pose parameters are determined using a more accurate method (e.g., the SLAM technique described herein) compared to one or more other joint poses or variables within the joint pose parameter set. In a specific example, the nonlinear solver may use the joint parameters and predetermined static weights to infer an upper body pose 205 that infers the most likely pose of one or more joints of user 102 at a given time or state.
[0047] While these techniques, used individually or in combination, generally allow for reliable determination of upper body posture 205 and related joints that are components of upper body posture 205 (e.g., head posture 210, wrist posture 220, elbow posture 230, etc.), these techniques may be unreliable and inaccurate for determining lower body posture 215 and related joints that are components of lower body posture 215 (e.g., hip posture 280, knee posture 290, ankle posture 295, etc.). Because user 102 only wears head-mounted device 104 and one or more controllers 106, the sensor data or image data associated with artificial reality system 100 may be limited. For example, image data received from one or more cameras 110 may not include at least a portion of user 102's lower body, resulting in limited information for determining or generating an accurate lower body posture 215 associated with user 102.
[0048] To address this issue, in a specific example, computing system 108 may utilize one or more techniques described herein to generate lower body pose 215 for user 102. In a specific example, lower body pose 215 may be generated by processing upper body pose 205 using a machine learning model. Figure 3 An example is shown where a lower body pose is generated using the input upper body pose. Specifically, Figure 3An upper-body pose 205, associated with user 102 of the artificial reality system 100, is shown. This upper-body pose 205 is generated using the methods described herein and received by a machine learning model 300. In a specific example, the machine learning model 300 may be based on a generative adversarial network (GAN). Using a specific example described herein, the machine learning model 300 may utilize the upper-body pose 205 to generate a lower-body pose 215. In a specific example, instead of or in addition to the lower-body pose 215 generated by the machine learning model, the computational system may combine the generated upper-body pose 205 with the generated lower-body pose 215 to generate a full-body pose.
[0049] Figure 4 A configuration for training a generative adversarial network (GAN) (also known as a machine learning model) 400 for pose prediction is shown. A GAN may include two separate neural networks: a generator 405 (which may be interchangeably referred to as "G" herein) and a discriminator 410 (which may be interchangeably referred to as "D" herein). In a particular example, the generator 405 and discriminator 410 may be implemented as, for example, but not limited to, neural networks, although any network architecture suitable for the operations described herein may be utilized. At a higher level, the generator 405 may be configured to receive a generated upper-body pose 205 as input and output a generated lower-body pose 215. In a particular example, the upper-body pose 205 may be combined with the lower-body pose 215 to generate a full-body pose 425. In a specific example, discriminator 410 may be configured to distinguish between "fake" full-body poses 425 (which include lower-body poses output by generator 405) and "true" trained full-body poses 435 from training pose database 440 (which are not generated by generator 405). In a specific example, one or more trained full-body poses 435 may include full-body poses from one or more images. Generator 405 and discriminator 410 can be considered adversaries, as generator 405 aims to generate fake poses that will fool discriminator 410 (in other words, to increase the error rate of discriminator 410), while discriminator 410 aims to correctly distinguish between "fake" poses from generator 405 and "true" poses. In a specific example, the machine learning model (e.g., generator 405) is optimized during training so that a second machine learning model (e.g., discriminator 410) incorrectly determines that a given full-body pose (which is generated using the machine learning model) is unlikely to have been generated using that machine learning model.
[0050] In a specific example, each trained full-body pose 435 in the training pose database can be a "true" pose in the sense that it is neither entirely nor partially generated by the machine learning model 400. In this specific example, each trained full-body pose 435 can be automatically obtained by retrieving different poses of a person from the training pose database, where each of these different poses includes the person's full-body pose 435. These "true" trained full-body poses 435 are used as ground truth to train the discriminator 410 network to identify which full-body poses are "true" and which are "false". The trained full-body poses 435 can also depict a person, such as, but not limited to, a person sitting, standing, or kneeling, or a person in another pose that is the same as or different from the generated upper-body pose 205 or the generated lower-body pose 215 output by the generator 405. The randomness of the poses makes the trained machine learning model 400 more robust because it may be unknown in operation what type of body pose the machine learning model 400 will be asked to process.
[0051] In a specific example, training of the machine learning model 400 can be performed simultaneously or in stages. For instance, a first stage can be used to train the generator 405, a second stage can be used to train the discriminator 410 based on the output of the generator 405, and a third stage can be used to retrain / refine the generator 405 to better "fool" the trained discriminator 410. As another example, training of the generator 405 and the discriminator 410 can occur simultaneously.
[0052] In a specific example, generator 405 may be configured to receive upper body pose 205 and generate a time-related lower body pose sequence 215, such as a sequence of lower body poses of a user standing, sitting, or walking. Instead of generating a single lower body pose 215, generator 405 may generate a time-related lower body pose sequence, which can be combined with corresponding upper body poses to generate a time-related full-body pose sequence. In a specific example, discriminator 410 may be configured and trained to determine whether the input pose sequence is temporally consistent, and thus distinguish between “false” time-related full-body pose sequences (including the time-related lower body pose sequences output by generator 405) and “true” time-related trained full-body pose sequences from training pose database 440 that were not generated by generator 405. In a specific example, the time-related trained full-body pose sequence may include full-body poses from one or more images. In this temporally correlated GAN, the generator 405 and discriminator 410 can be considered adversaries. The generator 405 aims to generate fake temporally correlated full-body pose sequences that will fool the discriminator 410 (in other words, to increase the discriminator 410's error rate), while the discriminator 410 aims to correctly distinguish between the "fake" and "real" temporally correlated full-body pose sequences from the generator 405. Once training is complete, the temporally correlated generator can be trained to output realistic temporally correlated lower-body pose sequences given an input upper-body pose 205. Although this disclosure primarily describes training and utilizing only one pose at a time using the method described herein for readability and clarity, alternatively, the GAN can be trained and used to output temporally correlated lower-body pose sequences using the same method described herein.
[0053] Figure 5 An example method 500 for training a generator based on loss is shown, based on a specific example. At a high level, during generator training 405, the parameters of generator 405 can be iteratively updated based on a comparison between the generated “fake” full-body pose 425 (which includes the lower-body pose 215 generated by the generator) and the corresponding training full-body pose 435. The goal of training is to maximize the prediction loss 460 of the “fake” pose. In doing so, the objective is for generator 405 to learn how to generate a lower-body pose 215 based on upper-body pose 205 such that the lower-body pose 215 can be used to generate a “fake” full-body pose 425 that looks “real” enough.
[0054] The method may begin at step 510, wherein the computing system may receive an upper body pose 205 corresponding to a first part of the user 102's body, the first part of which includes the user 102's head and arms. The upper body pose 205 input to generator 405 may be an upper body pose generated based on sensor data and / or image data using one or more techniques described herein; or the upper body pose 205 input to generator 405 may be an upper body pose extracted from one or more real images of a person.
[0055] At step 520, the machine learning model can process the upper body pose 205 to generate a lower body pose 215 corresponding to a second part of the user 102's body, which includes the user 102's legs. In a specific example, the generator 405 can generate the lower body pose 215 based on the received upper body pose 205. In a specific example, the generated lower body pose 215 corresponds to a second part of the body including the user's legs. The generated lower body pose 215 can be used to generate a "fake" full-body pose 425 that includes the generated lower body pose 215.
[0056] In a specific example, the computing system can determine contextual information associated with the artificial reality system 100. In a specific example, the generator can receive this contextual information. Generator 405 can utilize this contextual information to generate a lower body pose. This contextual information can be associated with a specific time when sensor data and / or image data is acquired. For example, and not limited to, the contextual information can include information about whether the user 102 is sitting or standing. In another example, the contextual information can include information about the application the user is interacting with at the time the upper body pose is determined, such as, but not limited to, artificial reality applications (e.g., whether the user is interacting with an application associated with a business meeting, or whether the user is interacting with an application associated with a game (e.g., a dance competition). In a specific example, generator 405 can receive one or more pieces of contextual information associated with the artificial reality system 100. In a specific example, the lower body pose 215 is generated by also processing the contextual information using a machine learning model.
[0057] In a specific example, the computational system may use, for example, physical perception data augmentation methods to determine one or more physical constraints associated with an upper-body or lower-body posture. Generator 405 may receive one or more of these physical constraints and utilize them to generate a more realistic lower-body posture or a more realistic time-related lower-body posture sequence. These physical constraints may be associated with one or more joint postures of the user, such as, but not limited to, physical restrictions on the possible range of motion of the knee joint posture, which simulates one or more physical limitations of the human body. In a specific example, the computational system may determine one or more physical constraints when generating a time-related lower-body posture sequence. For example, rather than restrictions, the computational system may determine time constraints on the rate of change of acceleration or jitter rate of one or more joints that are joints in one or more consecutive postures in a time-related posture sequence. In a specific example, generator 405 may use this constraint to determine a more realistic time-related lower-body posture sequence.
[0058] At step 530, the system can generate a full-body pose 425 based on the upper body pose 205 and the lower body pose 215. Figure 6 This illustrates the generation of a full-body pose from a generated upper-body pose and a generated lower-body pose. In a specific example, an upper-body pose 205 is generated using the methods described herein, and this upper-body pose 205 corresponds to a first part of the user's body, which includes the user's head and arms. In a specific example, a lower-body pose 215 is generated using the methods described herein, and this lower-body pose 215 corresponds to a second part of the user's body, which includes the user's legs. In a specific example, a full-body pose 425 can be generated by a machine learning model 400. In a specific example, the full-body pose 425 can be generated by a computational system in a post-processing step.
[0059] At step 540, the system can use discriminator 410 to determine whether the full-body pose 425 is likely generated using the lower-body pose 215 generated by generator 405 (i.e., whether the full-body pose 425 is "false" or "true"). In a particular example, the discriminator is trained to learn how to correctly distinguish between "true" and "false" full-body poses. For example, for each full-body pose, discriminator 410 can output a value between 0 and 1, representing the confidence or probability / likelihood that the pose was generated by the generator (i.e., "false"). For example, a value closer to 1 can indicate a higher probability / confidence that the pose is "true" (i.e., not generated by the generator), while a value closer to 0 can indicate a lower probability / confidence that the pose is "true" (which implicitly means a higher probability / confidence that the pose is "false"). These predictions can be compared to known labels for the poses, which indicate which poses are "false" and which are "true".
[0060] In a specific example, the system can calculate the loss based on D's judgment regarding whether the full-body posture could have been generated using the lower-body posture 215 generated by generator 405. For example, the loss can be calculated based on a comparison of the discriminator's prediction (e.g., confidence / probability score) with known labels indicating that the full-body posture is "false" (which could be implicit labels). In a specific example, for a given posture, the discriminator 410's prediction (represented as an output value between 0 and 1, indicating the confidence or probability / likelihood of the posture being generated by generator 405) can be compared to the posture's known labels, which indicate which postures are "false" and which are "true." For example, if the full-body posture is known to be "false," a prediction closer to 1 results in a higher loss, while a prediction closer to 0 results in a lower loss. In other words, the loss can be a measure of the correctness of the discriminator 410's predictions.
[0061] Then, at step 550, the system can update the generator based on this loss. The parameters of generator 410 can be updated to maximize the loss output by the discriminator. Since the generator's goal is to "fool" the discriminator into believing that the full-body pose generated using the lower-body pose 215 generated by the generator is "true," the training algorithm can be configured to optimize this loss. In other words, the generator's parameters are updated to optimize the generator to generate lower-body poses (which are subsequently used to generate full-body poses) that will increase the discriminator's prediction loss (i.e., increase the discriminator's incorrect prediction that a given full-body pose generated using the lower-body pose 215 generated by generator 405 is unlikely to have been generated using generator 405).
[0062] Then, at step 560, the system can determine whether training of generator 405 has been completed. In a specific example, the system can determine whether training is complete based on one or more termination criteria. For example, if the loss is below a predetermined threshold and / or if the loss has stabilized in the last few iterations (e.g., fluctuating within a predetermined range), the system can determine that training is complete. Alternatively or additionally, training can be considered complete if a predetermined number of training iterations have been completed or if a predetermined number of training samples have been used. In a specific example, if training is not yet complete, the system can repeat the process starting from step 510 to continue training the generator. Alternatively, if training is determined to be complete, the trained generator can be used in operation.
[0063] The training objective is to enable generator 405 to learn to generate poses 215 that will fool discriminator 410 into thinking they are "real". In a specific example, generator 405 can generate lower body poses 215, and discriminator 410 can predict the probability 450 that the resulting full-body pose 425 is "real" or "false". If a high prediction value indicates that discriminator 410 thinks the pose is more likely to be "real", then the training objective can be represented as maximizing discriminator 410's prediction of "false" full-body poses 425. If the discriminator 410's prediction accuracy for "false" full-body poses 425 is represented as a loss 460, then the training objective can be represented as maximizing that loss 460. Based on the loss function and the training objective, the parameters of generator 405 can be iteratively updated after each training iteration, making generator 405 better at generating lower body poses 215 that can "fool" discriminator 410. Therefore, once trained, generator 405 can be used to process a given upper body pose and automatically generate realistic lower body poses.
[0064] Return to Figure 4 According to a specific example, while generator 405 is being trained, it can also be used to simultaneously train discriminator 410. At a higher level, generator 405 can process a given upper body pose 205 and generate a lower body pose 215. The computational system can use the upper body pose 205 and the lower body pose 215 to generate a full-body pose 425. The generated full-body pose 425, along with one or more trained full-body poses 435 from the training pose database 440, can be provided as input to discriminator 410.
[0065] Figure 7An example method 700 for training the discriminator 410 is shown. In a particular example, the generator 405 can be used to train the discriminator 410. In a particular example, the generator 405 and the discriminator 410 can be trained simultaneously. The task of the discriminator 410 may be to process input poses (e.g., “false” full-body poses 425 and “true” training full-body poses 435) and predict which poses are “true” and which are “false” 450. Conceptually, the goal of training the discriminator 410 may be to maximize the predictions of the discriminator 410 for the training full-body pose 435 and minimize the predictions of the discriminator 410 for the generated full-body pose 425. In other words, if the correctness of the discriminator 410’s predictions for “false” poses (e.g., the generated full-body pose 425) and “true” poses (i.e., the training full-body pose 435) is represented by losses 460 and 465, respectively, then the goal of training the discriminator 410 would be to minimize these losses (i.e., minimize incorrect predictions). Based on the loss function and training objective, the parameters of the discriminator 410 can be iteratively updated after each prediction, making the discriminator 410 better able to distinguish between “true” and “false” full-body poses.
[0066] The method may begin at step 710, wherein the computing system may receive an upper body pose 205 corresponding to a first part of the user 102's body, the first part of which includes the user 102's head and arms. The upper body pose 205 input to generator 405 may be an upper body pose generated based on sensor data and / or image data using one or more techniques described herein; or the upper body pose 205 input to generator 405 may be an upper body pose extracted from one or more real images of a person. In a particular example, step 710 may be performed similarly to step 510 described herein.
[0067] At step 720, the machine learning model can process the upper body pose 205 to generate a lower body pose 215 corresponding to a second part of the user 102's body, including the user 102's legs. In a particular example, generator 405 can generate the lower body pose 215 based on the received upper body pose 205. In a particular example, the generated lower body pose 215 corresponds to a second part of the body including the user's legs. In a particular example, generator 405 can generate the lower body pose 215 based on its current parameters, which can be iteratively updated during training to allow the generator to generate more realistic lower body poses. The generated lower body pose 215 can be used to generate a "fake" full-body pose 425 including the generated lower body pose 215. In a particular example, step 720 can be performed similarly to step 520 described herein.
[0068] In a specific example, the computing system can determine contextual information associated with the artificial reality system 100. In a specific example, the generator can receive this contextual information. Generator 405 can utilize this contextual information to generate a lower body pose. This contextual information can be associated with a specific time when sensor data and / or image data is acquired. For example, and not limited to, the contextual information can include information about whether the user 102 is sitting or standing. In another example, the contextual information can include information about the application the user is interacting with at the time the upper body pose is determined, such as, but not limited to, an artificial reality application (e.g., whether the user is interacting with an application associated with a business meeting, or whether the user is interacting with an application associated with a game (e.g., a dance competition). In a specific example, generator 405 can receive one or more pieces of contextual information associated with the artificial reality system 100. In a specific example, the lower body pose 215 is generated by also processing the contextual information using a machine learning model.
[0069] At step 730, the system can generate a full-body pose 425 based on the upper-body pose 205 and the lower-body pose 215. In a specific example, the full-body pose 425 can be generated by a machine learning model 400. In a specific example, the full-body pose 425 can be generated by a computational system in a post-processing step. In a specific example, step 730 can be performed similarly to step 530 described herein.
[0070] At step 740, the system can use discriminator 410 to determine whether the full-body pose 425 could have been generated using a lower-body pose 215 that was generated using generator 405 (i.e., whether the full-body pose 425 is "false" or "true"). In a particular example, the discriminator is trained to learn how to correctly distinguish between "true" and "false" full-body poses. For example, for each full-body pose, discriminator 410 can output a value between 0 and 1, which represents the confidence or probability / likelihood that the pose was generated by the generator (i.e., "false"). For example, a value closer to 1 can indicate a higher probability / confidence that the pose is "true" (i.e., not generated by the generator), while a value closer to 0 can indicate a lower probability / confidence that the pose is "true" (which implicitly means a higher probability / confidence that the pose is "false"). These predictions can be compared to known labels for the poses, which indicate which poses are “false” and which are “true.” In a specific example, step 740 can be performed similarly to step 540 described herein.
[0071] In a specific example, the system can calculate the loss based on the discriminator's judgment regarding whether the full-body posture could have been generated using the lower-body posture 215 generated by generator 405. For example, the loss can be calculated based on a comparison of the discriminator's prediction (e.g., confidence / probability score) with known labels indicating that the full-body posture is "false" (this known label could be implicit). In a specific example, for a given posture, the discriminator 410's prediction (represented as an output value between 0 and 1, indicating the confidence or probability / likelihood of the posture being generated by generator 405) can be compared to the posture's known labels, which indicate which postures are "false" and which are "true." For example, if the full-body posture is known to be "false," a prediction closer to 1 might lead to a higher loss, while a prediction closer to 0 might lead to a lower loss. In other words, the loss can be a measure of the correctness of the discriminator 410's predictions.
[0072] Then, at step 750, the system can update the discriminator based on this loss. The parameters of the discriminator 410 can be updated to minimize the loss output by the discriminator. In a specific example, these losses can be backpropagated, and they can be used by a training algorithm to update the parameters of the discriminator 410, so that during training, the discriminator will gradually and better distinguish between "false" and "true" poses. The goal of the training algorithm can be to minimize this loss (i.e., minimize incorrect predictions).
[0073] Then, at step 760, the system can determine whether training of the discriminator 405 has been completed. In a specific example, the system can determine whether training is complete based on one or more termination criteria. For example, if the loss is below a predetermined threshold and / or if the loss has stabilized in the last few iterations (e.g., fluctuating within a predetermined range), the system can determine that training is complete. Alternatively or additionally, training can be considered complete if a predetermined number of training iterations have been completed or if a predetermined number of training samples have been used. In a specific example, if training is not yet complete, the system can repeat the process starting from step 710 to continue training the discriminator.
[0074] Although this disclosure will Figure 7 The method describes and illustrates multiple specific steps as occurring in a specific order, but this disclosure contemplates... Figure 7 Any suitable multiple steps in the method occur in any suitable order. Furthermore, although this disclosure describes and illustrates including... Figure 7 This disclosure describes example methods for training a discriminator that include these specific steps, but any suitable method for training a discriminator that includes any suitable steps, where appropriate, may include... Figure 7All steps, some steps, or none of the steps in the method Figure 7 Any step in the method. Furthermore, although this disclosure describes and illustrates the execution of... Figure 7 The method may involve specific components, devices, or systems for multiple specific steps, but this disclosure is contemplated for the execution of... Figure 7 Any suitable combination of any suitable component, device, or system in any suitable step of the method.
[0075] Return to Figure 5 Once generator 405 has been trained, it can be configured to receive an input upper-body pose 205 and generate a lower-body pose 215. After training, the lower-body pose will be realistic, such that when the lower-body pose is combined with the upper-body pose 205, the computational system will generate a realistic full-body pose 425. Furthermore, once trained, the generator can be used to generate lower-body poses based on any input upper-body pose (in other words, the generator is not limited to generating lower-body poses based on poses appearing in the training pose dataset). The trained generator can also be distributed to different platforms than the training system, including, for example, a user's mobile device or other personal computing device.
[0076] At step 570, the trained generator has access to upper body pose 205. Upper body pose 205 may correspond to a first part of the user 102's body, which includes the user 102's head and arms. The upper body pose 205 input to the trained generator may be an upper body pose generated based on sensor data using one or more techniques described herein, or the upper body pose 205 input to the trained generator 405 may be an upper body pose from one or more real images of a person.
[0077] Then, at step 580, the trained generator can generate a lower body pose corresponding to the second part of the user's body. The lower body pose can correspond to the second part of the user 102's body, which includes the user 102's legs. In a specific example, the trained generator can generate a lower body pose 215 based on the received upper body pose 205. In a specific example, the generated lower body pose 215 corresponds to the second part of the body, including the user's legs.
[0078] In a specific example, the trained generator can also receive and process contextual information that specifies additional information for generating an accurate lower body pose 215. In a specific example, the contextual information may be associated with the time at which sensor data used to generate the upper body pose 205 is acquired. For example, the generator may receive contextual information including an application that the user is interacting with when the sensor data is acquired, such as a virtual reality application for virtual business meetings. Based on this information, the trained generator 405 may be more likely to generate a lower body pose 215 that is typical of a user sitting or standing in a workplace environment.
[0079] Return to Figure 6 In a specific example, the computing system can generate a full-body pose 425 based on upper-body pose 205 and lower-body pose 215. In a specific example, the full-body pose 425 can be generated by a machine learning model 400. In a specific example, the full-body pose 425 can be generated by the computing system in a post-processing step. In a specific example, the full-body pose 425 can be used for various applications. For example, the computing system 108 can utilize the full-body pose 425 to generate an avatar of user 102 in a virtual reality space or an artificial reality space. In a specific example, the computing system 108 can utilize only a portion of the full-body pose 425 for various applications, which may be, for example, but is not limited to, an upper-body pose (e.g., from the user's head to the user's hips) or an inferred pose of only one or more joints (e.g., elbow 230 or knee 290).
[0080] Although this disclosure will Figure 5 The method describes and illustrates multiple specific steps as occurring in a specific order, but this disclosure contemplates... Figure 5 Any suitable multiple steps in the method occur in any suitable order. Furthermore, although this disclosure describes and illustrates example methods for training a generator based on encoded feature representations (the method includes...), Figure 5 (These specific steps in the method), however, this disclosure contemplates any suitable method for training a generator that includes any suitable steps, where appropriate, such suitable steps may include Figure 5 All steps, some steps, or none of the steps in the method Figure 5 Any step in the method. Furthermore, although this disclosure describes and illustrates the execution of... Figure 5 The method may involve specific components, devices, or systems for multiple specific steps, but this disclosure is contemplated for the execution of... Figure 5 Any suitable combination of any suitable component, device, or system in any suitable step of the method.
[0081] Although this disclosure will Figure 4 The method describes and illustrates multiple specific steps as occurring in a specific order, but this disclosure contemplates... Figure 4 Any suitable multiple steps in the method occur in any suitable order. Furthermore, although this disclosure describes and illustrates example methods for training a generator based on recurrence loss (the method includes...), Figure 4 These specific steps in the method), however, this disclosure contemplates any suitable method for training a generator based on reproducibility loss that includes any suitable steps, where appropriate, such suitable steps may include Figure 4 All steps, some steps, or none of the steps in the method Figure 4 Any step in the method. Furthermore, although this disclosure describes and illustrates the execution of... Figure 4 The method may involve specific components, devices, or systems for multiple specific steps, but this disclosure is contemplated for the execution of... Figure 4 Any suitable combination of any suitable component, device, or system in any suitable step of the method.
[0082] Figure 8 An example method 800 for generating a user's full-body posture based on upper and lower body postures is illustrated. The method may begin at step 810, where a computing system may receive sensor data acquired by one or more sensors associated with the user. In a particular example, the one or more sensors may be associated with user 102. In a particular example, the one or more sensors may be associated with a head-mounted device 104 worn by the user. For example, and not limitingly, head-mounted device 104 may include a gyroscope or inertial measurement unit that tracks the user's real-time movements, and head-mounted device 104 may output sensor data for representing or describing the movements. As another example, and not limitingly, the one or more controllers 106 may include an inertial measurement unit (IMU) and an infrared (IR) light-emitting diode (LED) configured to collect IMU sensor data and transmit the IMU sensor data to computing system 108. As another example, the sensor data may include one or more image data from one or more components of artificial reality system 100.
[0083] At step 820, the computing system can generate an upper-body pose corresponding to a first part of the user's body, which includes the user's head and arms. In a particular example, the computing system can utilize sensor data to generate the upper-body pose. In a particular example, the computing system 108 can determine multiple poses corresponding to multiple predetermined body parts of the user 102, such poses being, for example, but not limited to, joint poses (e.g., head pose 210 or wrist pose 220). In a particular example, a combination of one or more techniques can be used to determine or infer one or more joint poses that are components of the upper-body pose 200, such techniques being, for example, but not limited to, localization techniques (e.g., SLAM), machine learning techniques (e.g., neural networks), known spatial relationships with one or more artificial reality system components (e.g., known spatial relationships between head-mounted device 104 and head pose 210), visualization techniques (e.g., image segmentation), or optimization techniques (e.g., nonlinear solvers).
[0084] At step 830, the computational system can generate a lower-body pose corresponding to a second part of the user's body, including the user's legs. In a specific example, the computational system can utilize a machine learning model to generate the lower-body pose, such as, but not limited to, a generative adversarial network (GAN) including a generator network and a discriminator network. In a specific example, the generator can receive contextual information associated with the upper-body pose. The generator can utilize this contextual information to generate the lower-body pose. This contextual information can be associated with a specific time when sensor data and / or image data were acquired.
[0085] At step 840, the computing system can generate the user's full-body pose based on the upper-body and lower-body poses. In a specific example, the full-body pose can be generated using a machine learning model. In a specific example, the full-body pose can be generated by the computing system in a post-processing step. In a specific example, this full-body pose can be used for various applications. For example, the computing system can use the full-body pose to generate an avatar of the user in a virtual reality space or an artificial reality space.
[0086] Specific examples can be repeated where appropriate. Figure 8 One or more steps in the method. Although this disclosure will Figure 8 The method describes and illustrates multiple specific steps as occurring in a specific order, but this disclosure contemplates... Figure 8 Any suitable multiple steps in the method occur in any suitable order. Furthermore, although this disclosure describes and illustrates an example method for generating a user's full-body posture based on upper-body and lower-body postures (the method includes...), Figure 8(These specific steps in the method), however, this disclosure contemplates any suitable method for generating a user's full-body posture based on upper and lower body postures, including any suitable steps, and where appropriate, the method may include Figure 8 All steps, some steps, or none of the steps in the method Figure 8 Any step in the method. Furthermore, although this disclosure describes and illustrates the execution of... Figure 8 The method may involve specific components, devices, or systems for multiple specific steps, but this disclosure is contemplated for the execution of... Figure 8 Any suitable combination of any suitable component, device, or system in any suitable step of the method.
[0087] Figure 9 An example network environment 900 associated with a social networking system is shown. Network environment 900 includes a client system 930, a social networking system 960, and a third-party system 970, which are interconnected via a network 910. Although Figure 9 A specific arrangement of client system 930, social networking system 960, third-party system 970, and network 910 is shown, but this disclosure contemplates any suitable arrangement of client system 930, social networking system 960, third-party system 970, and network 910. By way of example and not limitation, two or more of client system 930, social networking system 960, and third-party system 970 may bypass network 910 and be directly connected to each other. As another example, two or more of client system 930, social networking system 960, and third-party system 970 may be physically or logically located in one place, wholly or partially. Furthermore, although... Figure 9 A specific number of client systems 930, social networking systems 960, third-party systems 970, and networks 910 are shown, but this disclosure contemplates any suitable number of client systems 930, social networking systems 960, third-party systems 970, and networks 910. As an example and not a limitation, network environment 900 may include multiple client systems 930, multiple social networking systems 960, multiple third-party systems 970, and multiple networks 910.
[0088] This disclosure considers any suitable network 910. By way of example and not limitation, one or more portions of network 910 may include an ad hoc network, intranet, extranet, virtual private network (VPN), local area network (LAN), wireless LAN (WLAN), wide area network (WAN), wireless LAN (WWAN), metropolitan area network (MAN), a portion of the Internet, a portion of the Public Switched Telephone Network (PSTN), a cellular telephone network, or a combination of two or more of these networks. Network 910 may include one or more networks 910.
[0089] Multiple links 950 can connect client system 930, social networking system 960, and third-party system 970 to network 910 or enable client system 930, social networking system 960, and third-party system 970 to connect to each other. This disclosure contemplates any suitable link 950. In a particular example, one or more links 950 include one or more wired (e.g., Digital Subscriber Line (DSL) or Data Over Cable Service Interface Specification (DOCSIS)) links, one or more wireless (e.g., Wi-Fi or Worldwide Interoperability for Microwave Access (WiMAX)) links, or one or more optical (e.g., Synchronous Optical Network (SONET) or Synchronous Digital Hierarchy (SDH)) links. In a specific example, one or more links 950 may each include an ad hoc network, intranet, extranet, VPN, LAN, WLAN, WAN, WWAN, MAN, a portion of the Internet, a portion of the PSTN, a cellular network, a satellite communication network, another link 950, or a combination of two or more such links 950. Throughout the network environment 900, the links 950 need not all be identical. One or more first links may differ from one or more second links in one or more respects.
[0090] In a specific example, client system 930 may be an electronic device comprising hardware, software, or embedded logic components, or a combination of two or more such components, and capable of performing the appropriate functions implemented or supported by client system 930. By way of example and not limitation, client system 930 may include a computer system, such as a desktop computer, notebook or laptop computer, netbook, tablet computer, e-book reader, GPS device, camera, personal digital assistant (PDA), handheld electronic device, cellular phone, smartphone, augmented / virtual reality device, other suitable electronic device, or any suitable combination thereof. This disclosure contemplates any suitable client system 930. Client system 930 enables network users at client system 930 to access network 910. Client system 930 enables its users to communicate with other users at other client systems 930.
[0091] In a specific example, client system 930 may include web browser 932 and may have one or more add-ons, plugins, or other extensions. A user at client system 930 may enter a Uniform Resource Locator (URL) or another address that directs web browser 932 to a specific server (e.g., server 962 or a server associated with third-party system 970), and web browser 932 may generate a Hypertext Transfer Protocol (HTTP) request and send that HTTP request to the server. The server may accept the HTTP request and, in response, send one or more Hypertext Markup Language (HTML) files to client system 930. Client system 930 may render the webpage based on the HTML file from the server for presentation to the user. This disclosure considers any suitable webpage file. By way of example and not limitation, a webpage may be rendered based on an HTML file, an Extensible Hypertext Markup Language (XHTML) file, or an Extensible Markup Language (XML) file, depending on specific needs. Such pages can also execute scripts, markup languages, and combinations of scripts. In this article, where appropriate, a reference to a webpage includes one or more corresponding webpage files (which the browser can use to render the webpage), and vice versa.
[0092] In a specific example, social networking system 960 may be a network-addressable computing system capable of controlling online social networks. Social networking system 960 can generate, store, receive, and send social network data, such as user profile data, concept profile data, social graph information, or other suitable data related to the online social network. Social networking system 960 may be accessed directly by other components in network environment 900 or via network 910. By way of example and not limitation, client system 930 may use a web browser 932 or a local application associated with social networking system 960 (e.g., a mobile social networking application, a messaging application, another suitable application, or any combination thereof) to access social networking system 960 directly or via network 910. In a specific example, social networking system 960 may include one or more servers 962. Each server 962 may be a single server or a distributed server spanning multiple computers or multiple data centers. Server 962 can be diverse, such as, but not limited to, a web server, news server, mail server, message server, advertising server, file server, application server, exchange server, database server, proxy server, another server suitable for performing the functions or processes described herein, or any combination thereof. In a particular example, each server 962 may include hardware, software, or embedded logic components, or combinations of two or more such components, for performing the appropriate functions implemented or supported by server 962. In a particular example, social networking system 960 may include one or more data storage areas 964. Data storage area 964 can be used to store various types of information. In a particular example, the information stored in data storage area 964 may be organized according to a particular data structure. In a particular example, each data storage area 964 may be a relational database, columnar database, correlation database, or other suitable database. Although this disclosure describes or illustrates a particular type of database, this disclosure contemplates any suitable type of database. Specific embodiments may provide multiple interfaces that enable client system 930, social network system 960, or third-party system 970 to manage, retrieve, modify, add, or delete information stored in data storage area 964.
[0093] In a specific example, the social network system 960 may store one or more social graphs in one or more data storage areas 964. In a specific example, the social graph may include multiple nodes—which may include multiple user nodes (each user node corresponds to a specific user) or multiple concept nodes (each concept node corresponds to a specific concept)—and multiple edges connecting these nodes. The social network system 960 may provide users of the online social network with the ability to communicate and interact with other users. In a specific example, a user can join an online social network via the social network system 960 and then add connections (e.g., relationships) to some other users in the social network system 960 that they wish to connect with. In this document, the term "friend" may refer to any other user with whom a user in the social network system 960 has already formed a connection, association, or relationship through the social network system 960.
[0094] In a specific example, the social networking system 960 may provide users with the ability to take action on various types of items or objects supported by the social networking system 960. By way of example and not limitation, these items and objects may include groups or social networks to which the user of the social networking system 960 may belong, events or calendar entries that the user may be interested in, computer-based applications that the user may use, transactions that allow the user to buy or sell items through the service, interactions with advertisements that the user may perform, or other suitable items or objects. Users may interact with anything that can be represented in the social networking system 960, or with anything that can be represented by an external system 970, which is separate from the social networking system 960 and coupled to the social networking system 960 via network 910.
[0095] In a specific example, the social networking system 960 may be able to connect various entities. By way of example and not limitation, the social networking system 960 may enable users to interact with each other and receive content from third-party systems 970 or other entities, or allow users to interact with these entities through application programming interfaces (APIs) or other communication channels.
[0096] In a specific example, third-party system 970 may include one or more types of servers, one or more data stores, one or more interfaces (including but not limited to APIs), one or more web services, one or more content sources, one or more networks, or any other suitable components (e.g., servers may communicate with these components). Third-party system 970 may be operated by an entity different from the entity operating social networking system 960. However, in a specific example, social networking system 960 and third-party system 970 may work together to provide social networking services to users of social networking system 960 or to third-party system 970. In this sense, social networking system 960 may provide a platform or backbone network that other systems (e.g., third-party system 970) can use to provide social networking services and functionality to users on the Internet.
[0097] In a specific example, third-party system 970 may include a third-party content object provider. The third-party content object provider may include one or more content object sources that can be transmitted to client system 930. As an example, and not a limitation, content objects may include information about things or activities that the user is interested in, such as movie showtimes, movie reviews, restaurant reviews, restaurant menus, product information and reviews, or other suitable information. As another example, and not a limitation, content objects may include incentivized content objects, such as coupons, discount vouchers, gift certificates, or other suitable incentivized content objects.
[0098] In a specific example, the social networking system 960 also includes user-generated content objects, which can enhance user interaction with the social networking system 960. User-generated content can include any content that a user can add, upload, send, or "post" to the social networking system 960. As an example, and not a limitation, a user transmits a post from the client system 930 to the social networking system 960. Posts can include data such as status updates or other text data, location information, photos, videos, links, music, or other similar data or media. Content can also be added to the social networking system 960 by third parties through "communication channels" such as news feeds or streams.
[0099] In a specific example, the social networking system 960 may include various servers, subsystems, programs, modules, logs, and data storage areas. In a specific example, the social networking system 960 may include one or more of the following: a web server, an action logger, an API request server, a relevance and ranking engine, a content object classifier, a notification controller, action logs, third-party content object exposure logs, an inference module, an authorization / privacy server, a search module, an ad targeting module, a user interface module, a user profile store, a contact store, a third-party content store, or a location store. The social networking system 960 may also include suitable components, such as a web interface, security mechanisms, a load balancer, a failover server, an administration and network operations console, other suitable components, or any suitable combination thereof. In a specific example, the social networking system 960 may include one or more user profile stores for storing user profiles. User profiles may include, for example, biometric information, demographic information, behavioral information, social information, or other types of descriptive information (e.g., work experience, educational history, hobbies or preferences, interests, close relationships, or location). Interest information may include interests associated with one or more categories. These categories can be generic or specific. As an example, and not a limitation, if a user “likes” an article about a shoe brand, the category could be that brand, or the generic category “shoes” or “clothing.” A contact store can be used to store contact information about users. Contact information can indicate users who have similar or shared work experience, group memberships, hobbies, educational history, or are in any way related to or share common attributes. Contact information can also include user-defined connections between different users and content (both internal and external). A web server can be used to link the social networking system 960 to one or more client systems 930 or one or more third-party systems 970 via network 910. The web server can include a mail server or other messaging functionality for receiving and sending messages between the social networking system 960 and one or more client systems 930. An API request server can allow third-party systems 970 to access information from the social networking system 960 by calling one or more APIs. An action logger can be used to receive communications from the web server regarding user actions on or outside the social networking system 960. Combined with the action log, a log of third-party content objects that a user is exposed to can be maintained. The notification controller can provide information about content objects to the client system 930. Information can be pushed to the client system 930 as a notification, or information can be retrieved from the client system 930 in response to a received request.An authorization server can be used to enforce one or more privacy settings for users of the social networking system 960. A user's privacy settings determine how specific information associated with that user can be shared. The authorization server can allow users, for example, by setting appropriate privacy settings, to choose whether or not their actions are recorded by the social networking system 960 or shared with other systems (e.g., third-party system 970). A third-party content object store can be used to store content objects received from third parties (e.g., third-party system 970). A location store can be used to store location information received from a client system 930 associated with the user. An advertising targeting module can combine social information, current time, location information, or other suitable information to deliver relevant advertisements to users in the form of notifications.
[0100] Figure 10 An example computer system 1000 is illustrated. In a particular example, one or more computer systems 1000 perform one or more steps of one or more methods described or illustrated herein. In a particular example, one or more computer systems 1000 provide the functionality described or illustrated herein. In a particular example, software running on one or more computer systems 1000 performs one or more steps of one or more methods described or illustrated herein, or provides the functionality described or illustrated herein. The particular example includes one or more portions of one or more computer systems 1000. In this document, references to computer systems may include computing devices and vice versa, where appropriate. Furthermore, references to computer systems may include one or more computer systems, where appropriate.
[0101] This disclosure contemplates any suitable number of computer systems 1000. This disclosure contemplates computer systems 1000 employing any suitable physical form. By way of example and not limitation, computer system 1000 may be an embedded computer system, a system-on-chip (SOC), a single-board computer system (SBC) (e.g., a computer-on-module (COM) or system-on-module (SOM)), a desktop computer system, a laptop or notebook computer system, an interactive self-service machine, a mainframe, a network of computer systems, a mobile phone, a personal digital assistant (PDA), a server, a tablet computer system, an augmented / virtual reality device, or a combination of two or more of these computer systems. Where appropriate, computer system 1000 may include one or more computer systems 1000; computer system 1000 may be single or distributed; spanning multiple locations; spanning multiple machines; spanning multiple data centers; or located in the cloud (which may include one or more cloud components in one or more networks). Where appropriate, one or more computer systems 1000 can perform one or more steps of the methods described or illustrated herein without significant space or time constraints. By way of example and not limitation, one or more computer systems 1000 can perform one or more steps of the methods described or illustrated herein in real time or in batch mode. Where appropriate, one or more computer systems 1000 can perform one or more steps of the methods described or illustrated herein at different times or in different locations.
[0102] Computer system 1000 includes a processor 1002, memory 1004, storage device 1006, input / output (I / O) interface 1008, communication interface 1010, and bus 1012. Although this disclosure describes and illustrates a particular computer system having a particular number of components in a particular arrangement, this disclosure contemplates any suitable computer system having any suitable number of components in any suitable arrangement.
[0103] In a particular example, processor 1002 includes hardware for executing a plurality of instructions, such as those that constitute a computer program. By way of example, and not limitation, to execute the plurality of instructions, processor 1002 may retrieve (or read) these instructions from internal registers, internal cache, memory 1004, or storage device 1006; decode and execute these instructions; and then write one or more results to the internal registers, internal cache, memory 1004, or storage device 1006. In a particular example, processor 1002 may include one or more internal caches for data, instructions, or addresses. Where appropriate, this disclosure contemplates processor 1002 including any suitable number of suitable internal caches. By way of example, and not limitation, processor 1002 may include one or more instruction caches, one or more data caches, and one or more translation lookaside buffers (TLBs). The plurality of instructions in the instruction cache may be copies of the plurality of instructions in memory 1004 or storage device 1006, and the instruction cache may accelerate the retrieval of those instructions by processor 1002. The data in the data cache may be a copy of the data in memory 1004 or storage device 1006 for operation by instructions executed at processor 1002; the result of a previous instruction executed at processor 1002 for access by subsequent instructions executed at processor 1002, or for writing to memory 1004 or storage device 1006; or the data in the data cache may be other suitable data. The data cache can accelerate read or write operations of processor 1002. Multiple TLBs can accelerate virtual address translation of processor 1002. In a particular example, processor 1002 may include one or more internal registers for data, instructions, or addresses. Where appropriate, this disclosure contemplates processor 1002 including any suitable number of suitable internal registers. Where appropriate, processor 1002 may include one or more arithmetic logic units (ALUs); processor 1002 may be a multi-core processor, or may include one or more processors 1002. Although this disclosure describes and illustrates specific processors, this disclosure contemplates any suitable processor.
[0104] In a particular example, memory 1004 includes main memory for storing instructions to be executed by processor 1002 or data to be operated by processor 1002. By way of example and not limitation, computer system 1000 may load multiple instructions from storage device 1006 or another source (e.g., another computer system 1000) into memory 1004. Processor 1002 may then load these instructions from memory 1004 into internal registers or internal cache memory. To execute these instructions, processor 1002 may retrieve and decode these instructions from internal registers or internal cache memory. During or after the execution of these instructions, processor 1002 may write one or more results (which may be intermediate or final results) into internal registers or internal cache memory. Processor 1002 may then write one or more of those results into memory 1004. In a particular example, processor 1002 executes only instructions in one or more internal registers or one or more internal caches, or in memory 1004 (different from memory device 1006 or other locations), and operates only on data in one or more internal registers or one or more internal caches, or in memory 1004 (different from memory device 1006 or other locations). One or more memory buses (each memory bus may include an address bus and a data bus) couple processor 1002 to memory 1004. As described below, bus 1012 may include one or more memory buses. In a particular example, one or more memory management units (MMUs) are located between processor 1002 and memory 1004 and facilitate access to memory 1004 requested by processor 1002. In a particular example, memory 1004 includes random access memory (RAM). Where appropriate, the RAM may be volatile memory. Where appropriate, the RAM may be dynamic RAM (DRAM) or static RAM (SRAM). Furthermore, where appropriate, the RAM may be a single-port RAM or a multi-port RAM. This disclosure contemplates any suitable RAM. Where appropriate, memory 1004 may include one or more memories 1004. Although this disclosure describes and illustrates specific memories, this disclosure contemplates any suitable memories.
[0105] In a particular example, storage device 1006 includes a mass storage device for data or instructions. By way of example, and not limitation, storage device 1006 may include a hard disk drive (HDD), a floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these storage devices. Where appropriate, storage device 1006 may include removable or non-removable (or fixed) media. Where appropriate, storage device 1006 may be internal or external to computer system 1000. In a particular example, storage device 1006 is a non-volatile solid-state memory. In a particular example, storage device 1006 includes read-only memory (ROM). Where appropriate, the ROM may be a mask-programmable ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM), or flash memory, or a combination of two or more of these ROMs. This disclosure contemplates a high-capacity storage device 1006 in any suitable physical form. Where appropriate, storage device 1006 may include one or more storage control units facilitating communication between processor 1002 and storage device 1006. Where appropriate, storage device 1006 may include one or more storage devices 1006. Although this disclosure describes and illustrates specific storage devices, this disclosure contemplates any suitable storage device.
[0106] In a particular example, I / O interface 1008 includes hardware, software, or both hardware and software that provide one or more interfaces for communication between computer system 1000 and one or more I / O devices. Where appropriate, computer system 1000 may include one or more of these I / O devices. These one or more I / O devices enable communication between a person and computer system 1000. By way of example and not limitation, I / O devices may include a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, still camera, stylus, input pad, touchscreen, trackball, camera, another suitable I / O device, or a combination of two or more of these I / O devices. I / O devices may include one or more sensors. This disclosure contemplates any suitable I / O device and any suitable I / O interface 1008 for such I / O devices. Where appropriate, I / O interface 1008 may include one or more device or software drivers that enable processor 1002 to drive one or more of these I / O devices. Where appropriate, I / O interface 1008 may include one or more I / O interfaces 1008. Although this disclosure describes and illustrates specific I / O interfaces, this disclosure considers any suitable I / O interface.
[0107] In a particular example, communication interface 1010 includes hardware, software, or both hardware and software providing one or more interfaces for communication (e.g., packet-based communication) between computer system 1000 and one or more other computer systems 1000 or with one or more networks. By way of example, and not limitation, communication interface 1010 may include a network interface controller (NIC) or network adapter for communicating with Ethernet or other wire-based networks, or a wireless NIC (WNIC) or wireless adapter for communicating with wireless networks such as Wi-Fi networks. This disclosure contemplates any suitable network and any suitable communication interface 1010 for that network. By way of example, and not limitation, computer system 1000 may communicate with one or more portions of an ad hoc network, personal area network (PAN), local area network (LAN), wide area network (WAN), metropolitan area network (MAN), or the Internet, or a combination of two or more of these networks. One or more of these networks may be wired or wireless. As an example, computer system 1000 may communicate with a wireless PAN (WPAN) (e.g., Bluetooth WPAN), a Wi-Fi network, a Wi-Fi Max network, a cellular telephone network (e.g., a Global System for Mobile Communication (GSM) network), or other suitable wireless networks, or a combination of two or more of these networks. Where appropriate, computer system 1000 may include any suitable communication interface 1010 for any of these networks. Where appropriate, communication interface 1010 may include one or more communication interfaces 1010. Although specific communication interfaces are described and illustrated in this disclosure, any suitable communication interface is contemplated herein.
[0108] In a particular example, bus 1012 includes hardware, software, or both hardware and software that couple multiple components of computer system 1000 to each other. By way of example and not limitation, bus 1012 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a front-side bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infiniband interconnect, a low-pin-count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCIe) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association local (VLB) bus, or another suitable bus, or a combination of two or more of these buses. Where appropriate, bus 1012 may include one or more buses 1012. Although this disclosure describes and illustrates a particular bus, this disclosure considers any suitable bus or interconnect.
[0109] In this document, where appropriate, a computer-readable non-transitory storage medium may include one or more semiconductor-based integrated circuits (ICs) or other integrated circuits (e.g., field-programmable gate arrays (FPGAs) or application-specific ICs (ASICs)), hard disk drives (HDDs), hybrid hard drives (HHDs), optical discs, optical disc drives (ODDs), magneto-optical disks, magneto-optical disk drives, floppy disks, floppy disk drives (FDDs), magnetic tape, solid-state drives (SSDs), RAM drives, secure digital cards, secure digital drives, any other suitable computer-readable non-transitory storage media, or any suitable combination of two or more of these storage media. Where appropriate, a computer-readable non-transitory storage medium may be volatile, non-volatile, or a combination of volatile and non-volatile computer-readable non-transitory storage media.
[0110] In this document, unless otherwise expressly indicated or the context otherwise indicates, “or” is inclusive rather than exclusive. Therefore, in this document, unless otherwise expressly indicated or the context otherwise indicates, “A or B” means “A, B, or both A and B”. Furthermore, unless otherwise expressly indicated or the context otherwise indicates, “and” is both common and separate. Therefore, in this document, unless otherwise expressly indicated or the context otherwise indicates, “A and B” means “A and B, commonly or separately”.
[0111] The scope of this disclosure covers all changes, substitutions, variations, alterations, and modifications to the examples and embodiments described or illustrated herein that will be understood by those skilled in the art. The scope of this disclosure is not limited to the examples described or illustrated herein. Furthermore, although this disclosure describes and illustrates various examples herein as including specific components, elements, features, functions, operations, or steps, those skilled in the art will understand that any example of these examples may include any combination or arrangement of any component, element, feature, function, operation, or step described or illustrated anywhere herein. Moreover, references in the appended claims to apparatus or systems, or components in apparatus or systems (that are adapted, arranged, capable, configured, implemented, operable, or usable to perform a particular function) cover that apparatus, system, or component (whether or not the apparatus, system, component, or the particular function is activated, turned on, or unlocked), provided that the apparatus, system, or component is so adapted, arranged, capable, configured, implemented, operable, or usable. Furthermore, although this disclosure describes or illustrates specific examples to provide specific advantages, specific examples may not provide these advantages, or may provide some or all of these advantages.
Claims
1. A method comprising: by a computing system: receiving sensor data collected by one or more sensors coupled to a user; determining contextual information associated with a time at which the sensor data was collected; generating, based on the sensor data, an upper body pose corresponding to a first portion of a body of the user, the first portion of the body including a head and arms of the user; generating, by processing the upper body pose using a machine learning model, a lower body pose corresponding to a second portion of the body of the user, the second portion of the body including legs of the user, wherein the lower body pose is generated by also processing the contextual information using the machine learning model; and generating, based on the upper body pose and the lower body pose, a full body pose of the user.
2. The method of claim 1, wherein, the machine learning model is trained using a second machine learning model trained to determine whether a given full body pose is likely to have been generated using the machine learning model.
3. The method of claim 1 or 2, wherein, the machine learning model is optimized during training such that the second machine learning model erroneously determines that a given full body pose generated using the machine learning model is less likely to have been generated using the machine learning model.
4. The method of claim 1 or 2, wherein, the machine learning model is trained by: generating a second lower body pose by processing a second upper body pose using the machine learning model; generating a second full body pose based on the second upper body pose and the second lower body pose; using a second machine learning model to determine whether the second full body pose is likely to have been generated using the machine learning model; and updating the machine learning model based on the determination of the second machine learning model. the one or more sensors are associated with a head-mounted device worn by the user.
5. The method of claim 1 or 2, wherein, generating the upper body pose includes:
6. The method of claim 1 or 2, wherein, determining, based on the sensor data, a plurality of poses corresponding to a plurality of predetermined body parts of the user; and inferring the upper body pose based on the plurality of poses. the plurality of poses includes a wrist pose corresponding to a wrist of the user, wherein the wrist pose is determined based on one or more images captured by a head-mounted device worn by the user, the one or more images depicting the wrist of the user or a device held by the user.
7. The method of claim 6, wherein, the contextual information includes an application with which the user is interacting at a time of determining the upper body pose.
8. The method of claim 1, wherein, generating an avatar of the user based on the full body pose.
9. The method of claim 1 or 2, further comprising:
10. One or more computer-readable non-transitory storage media comprising instructions to, when executed by a server, cause the server to perform the method of any of the preceding claims.
11. A system comprising: one or more processors; and the one or more computer-readable non-transitory storage media of claim 10 coupled to the one or more processors.
Citation Information
Patent Citations
Joint infrared and visible light visual-inertial object tracking
US20210208673A1
Methods and devices for assessing a captured motion
US20180070864A1
Systems and methods for generating complementary data for visual display
WO2020060666A1