System and method for human pose and shape estimation
The artificial neural network system generates a human body model based on joint position information, which solves the problem of difficulty in real-time restoring the patient's posture and shape in medical settings, and achieves higher positioning and navigation accuracy.
Patent Information
- Application Number
- CN202011352135.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-08-17
- Filing Date
- 2020-11-26
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2040-11-26
AI Technical Summary
The prior art is difficult to restore the patient's posture and shape in real time in medical settings, especially if the patient is blocked by medical equipment or clothing.
An artificial neural network (ANN) system is used to generate a human body model based on joint position information, including posture and shape parameters. The system can use joint position information in the training data to predict and restore the patient's posture and shape by optimizing execution parameters.
The ability to restore patient posture and shape under limited information conditions is achieved, improving the accuracy of patient positioning and surgical navigation in medical settings.
Smart Images

Figure CN112419419B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of Provisional U.S. Patent Application No. 62 / 941,203 filed on November 27, 2019 and Provisional U.S. Patent Application No. 16 / 995,446 filed on August 17, 2020, the disclosures of which are incorporated herein by reference in their entirety. Technical Field
[0003] The present application relates to the field of human body model estimation. Background Art
[0004] Computer-generated human models that truly represent the posture and shape of a patient can be used in a wide range of medical applications, including, for example, patient positioning, surgical navigation, unified medical record analysis, and the like. For example, in radiation therapy and medical imaging, success often depends on the ability to place and maintain the patient in a desired position so that treatment or scanning can be performed in a precise and accurate manner. Having real-time knowledge of the patient's physical characteristics in these situations (such as the patient's body shape and posture) can bring many benefits, such as positioning the patient faster and more accurately according to a scan or treatment plan, obtaining more consistent results in multiple scans or multiple treatment sessions, and the like. In other example cases, such as during a surgical procedure, information about the patient's body shape can provide insight and guidance for both surgical planning and execution. For example, the information can be used to locate and navigate around the patient's surgical area. And when presented visually in real time, the information can also provide a device for monitoring the patient's status during the surgical procedure.
[0005] Conventional techniques for restoring a patient's body model rely on having comprehensive knowledge about the patient's joint positions, and may only be able to restore the patient's posture based on the joint positions. However, in many medical settings, the patient's body is often at least partially obscured by medical equipment and / or clothing, and obtaining information about the patient's body shape in these settings may also be beneficial. Therefore, it is highly desirable to have the ability to restore both the patient's posture and shape, even with only limited information about the patient's physical features. Summary of the invention
[0006] This article describes systems, methods, and devices for recovering a person's posture and shape based on one or more joint positions of the person. The joint positions can be determined based on an image of the person, such as an image including color and / or depth information representing physical features of the person. The joint positions can be a subset of all joint positions of the person (e.g., excluding joints that are occluded or otherwise unknown). The system, method, and device may include one or more processors and one or more storage devices storing instructions, which, when executed by one or more processors, cause one or more processors to estimate with an artificial neural network and provide the artificial neural network with information related to one or more joint positions of the person. The artificial neural network can determine a first plurality of parameters associated with the person's posture and a second plurality of parameters associated with the person's shape based on information related to one or more joint positions of the person. Based on the first plurality of parameters and the second plurality of parameters, one or more human models representing the person's posture and shape can be generated.
[0007] The artificial neural network can be trained using training data to perform one or more of the above-mentioned tasks, and the training data includes the joint positions of the human body. During training, the artificial neural network can predict the posture and shape parameters associated with the human body based on the joint positions included in the training data. The artificial neural network can then infer the joint positions of the human body from the predicted posture and shape parameters, and adjust (e.g., optimize) the execution parameters (e.g., weights) of the artificial neural network based on the difference between the inferred joint positions and the joint positions included in the training data. In an example, the training data may also include posture and shape parameters associated with the joint positions of the human body, and the artificial neural network may also adjust (e.g., optimize) its execution parameters based on the difference between the predicted posture and shape parameters and the posture and shape parameters included in the training data.
[0008] In order to acquire the ability to predict pose and shape parameters based on partial knowledge about the joint positions of a person (e.g., some joint positions of a person may be occluded, unobservable, or otherwise unknown to the neural network), training the artificial neural network may involve providing a subset of joint positions to the artificial neural network and forcing the artificial neural network to use the subset of joint positions to predict pose and shape parameters. For example, the training may utilize an existing parametric human body model associated with the human body to determine multiple joint positions of the human body, and then randomly exclude a subset of multiple joint positions from the input of the artificial neural network (e.g., by artificially treating the subset of joint positions as unobserved and unavailable).
[0009] The joint positions described herein may include two-dimensional (2D) and / or three-dimensional (3D) joint positions of a person. When at least 2D joint positions are used for training, the artificial neural network may predict posture and shape parameters based on the 2D joint positions during training, infer the 3D joint positions of the human body using the predicted posture and shape parameters, and project the 3D joint positions into the image space to obtain corresponding 2D joint positions. The artificial neural network may then adjust its execution parameters based on the difference between the projected 2D joint positions and the 2D joint positions included in the training data.
[0010] The pose and shape parameters described herein can be recovered separately (e.g., independently of each other). For example, the recovered pose parameters can be independent of body shape (e.g., independent of a person's height and weight). Therefore, in addition to medical applications, the techniques described herein can be used together with various human motion capture systems (e.g., systems that can output 3D joint positions) in, for example, game development, animation development, special effects for movies, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Examples disclosed herein may be understood in more detail from the following description given by way of example in conjunction with the accompanying drawings.
[0012] Figure 1 is a block diagram illustrating an example system for recovering the pose and shape of a person based on joint position information associated with the person.
[0013] Figure 2 is a block diagram illustrating an example of training an artificial neural network to learn a model for predicting a person's posture and shape.
[0014] Figure 3 is a flow chart illustrating an example neural network training process.
[0015] Figure 4 is a block diagram illustrating an example neural network system described herein. DETAILED DESCRIPTION
[0016] The present disclosure is illustrated by way of example and not limitation in the figures of the accompanying drawings.
[0017] Figure 11 is a diagram illustrating an example system 100, which is configured to restore the posture and shape of a person, and generates a human body model representing the posture and shape of such a person based on the joint position information associated with the person. As shown in the figure, the system 100 can be configured to receive information about the joint position 102 of the person. The joint position 102 may include a 2D joint position (e.g., a 2D coordinate representing the 2D joint position, a 2D key point, a 2D feature (binary heat map), etc.). The joint position 102 may include a 3D joint position (e.g., a 3D coordinate representing the 3D joint position, a 3D key point, a 3D feature map, etc.). The joint position 102 may also include a combination of a 2D joint position and a 3D joint position. The joint position 102 can be derived from an image of the person, such as a color image of the person (e.g., a red-green-blue or RGB image), a depth image of the person, or an image of the person including both color information and depth information (e.g., a color plus depth or RGB-D image). Such an image can be captured by a sensor (e.g., an RGB sensor, a depth sensor, an infrared sensor, etc.), a camera, or another suitable device (e.g., a medical imaging device) capable of producing a visual representation of a person with a specific posture and body shape. The 2D joint positions and / or 3D joint positions can be derived, for example, based on 2D features and / or 3D features extracted from the image using one or more convolutional encoders. Examples of convolutional encoders can be found in commonly assigned U.S. patent application Ser. No. 16 / 863,382, filed on April 30, 2020, entitled “Systems and Methods for Human Mesh Recovery,” the disclosure of which is hereby incorporated by reference in its entirety.
[0018] The joint positions 102 may also be derived from a system or application configured to output joint position information of a person. Such a system or application may be, for example, a human motion capture system used in game development, animation development, special effects generation (e.g., for movies), etc.
[0019] The joint positions 102 may not include all joint positions of the person as depicted in the image or output by the human motion capture system. For example, one or more joint positions of the person may be occluded or unobservable in the image, mispredicted by an upstream device or program, or otherwise unknown to the system 100. Therefore, the joint positions 102 used by the system 100 to predict the pose and shape of the person may include only a subset of the joint positions of the person.
[0020] The system 100 may include a posture and shape regressor 104, which is configured to receive joint positions 102 (e.g., as input) and determine multiple parameters θ associated with the posture of the person and multiple parameters β associated with the shape of the person based on the joint positions 102. The posture and shape regressor 104 can be implemented using an artificial neural network (ANN), which includes multiple layers, such as one or more input layers, one or more hidden layers, and / or one or more output layers. The input layer of the ANN can be configured to receive the joint positions 102 and pass them to subsequent layers for processing. Each input layer can include one or more channels, and each channel can be configured to receive data from a corresponding data source. The hidden layer of the ANN may include one or more convolutional layers, one or more pooling layers, and / or one or more fully connected (FC) layers (e.g., regression layers). For example, the hidden layer may include multiple (e.g., a stack of ten) FC layers with a rectified linear unit (ReLU) activation function, and each FC layer may include multiple units (e.g., neurons) with corresponding weights that, when applied to features associated with the joint position 102, regress posture parameters θ and shape parameters β based on the features associated with the joint position 102.
[0021] In an example, the N joints can be represented by the following vectors: [x1, y1, x2, y2, x3, y3...], where x and y represent the positions of each joint. This can produce a 2N-dimensional vector corresponding to x and y for each of the N joints. In response to receiving such a vector input, the ANN (e.g., the FC layer of the ANN) can progressively transform the input vector into a different dimension. Each FC unit can be followed by a nonlinear activation unit (e.g., ReLU) and / or the next FC unit. Based on this operation, an output vector with K dimensions can be derived, where K can represent the number of estimated parameters. For example, if the posture parameter θ has 75 dimensions and the shape parameter β has 10 dimensions, then K can be equal to 85 (e.g., K=75+10).
[0022] In the example, the joint positions can be represented by feature maps such as binary heat maps, from which it can be seen that the joints can have a matrix representation. Multiple such matrices (e.g., corresponding to N joints) can be combined to produce N-channel (e.g., one channel per joint) inputs to the ANN. Before the FC unit of the ANN reaches the apex, the input can be processed by one or more convolutional layers of the ANN to produce a K-dimensional output vector representing the posture parameters θ and the shape parameters β.
[0023] In an example, the regressed posture parameter θ may include 72 parameters (e.g., 3 parameters for each of the 23 joints and 3 parameters for the root joint), and the regressed shape parameter β may include multiple coefficients (e.g., the first 10 coefficients) of a principal component analysis (PCA) space. Once the posture and shape parameters are determined, a plurality of vertices (e.g., 6890 vertices based on 72 posture parameters and 10 shape parameters) may be obtained for constructing a human body model representing the posture and shape of a person. For example, the human body model may be a statistical parameter differential model, such as a skinned multi-person linear (SMPL) model that defines the following function: M(β; θ; Φ): (where Φ may represent learned SMPL model parameters), which is used to generate N mesh vertices associated with a 3D mesh of a person. The mesh vertices may be used by a mesh generator 108 of the system 100 to generate a visual representation 110 of a human body model, for example, by shaping template body vertices conditioned on θ and β, articulating bones according to joint rotations indicated by θ (e.g., via forward kinematics), and deforming the 3D surface using linear blend skinning.
[0024] The pose and shape regressor 104 can be configured to recover the pose parameters θ separately or independently from the shape parameters β. This separation or independence can be achieved, for example, by applying pose normalization during the recovery operation, so that the pose parameters can be estimated without knowledge of the subject's specific body shape, and vice versa. The pose parameters θ thus recovered can be independent of body shape (e.g., independent of the person's height and / or weight), and similarly, the shape parameters β recovered can also be independent of pose (e.g., independent of the person's joint angles). In this way, one of the recovered pose parameters or shape parameters can be used in a system or application (e.g., a human motion capture system) without the other of the recovered pose parameters or shape parameters.
[0025] The pose and shape regressor 104 can be trained to predict pose parameters θ and shape parameters β using training data including joint position information of a human body. The joint position information can be paired with human body model parameters (e.g., SMPL parameters) to introduce a cycle consistency objective into the training. Figure 2204 (e.g., pose and shape regressor 104) to learn a model for predicting pose and shape parameters associated with a human body model based on joint position information. Training can be performed using data including joint position information 202 (e.g., 2D or 3D key points or features representing joint positions) and / or corresponding parameters 206 of a human body model from which joint positions 202 can be derived. Such paired training data can be obtained, for example, from a publicly available human motion capture (MoCap) database and used as real data for optimizing the parameters of the ANN 204.
[0026] During training, the ANN 204 may receive joint position information 202 and predict pose and shape parameters 208 based on the joint position information using initial execution parameters of the ANN 204 (e.g., initial weights associated with one or more fully connected layers of the ANN 204). The initial execution parameters may be derived, for example, by sampling them from one or more probability distributions or based on parameter values of another neural network with a similar architecture. When making predictions, the ANN 204 may compare the predicted pose and shape parameters 208 with the true data parameters 206 and calculate parameter losses 210 according to a loss function. The loss function may be, for example, based on the Euclidean distance between the predicted pose and shape parameters 208 and the true data parameters 206, for example, in, represents the predicted pose and shape parameters 208, and [β,θ] represents the real data pose and shape parameters 206. The loss function can also be based on L1 distance, Haussdorf distance, etc.
[0027] ANN 204 can additionally consider joint position loss when adjusting its execution parameters. For example, the training of ANN 204 can also utilize joint position regressor 212 to infer joint position 214 (for example, 3D joint position) based on predicted posture and shape parameters 208, and adjust the execution parameters of ANN 204 by minimizing the difference between the inferred joint position 214 and the input joint position 202. Joint position regressor 212 (for example, it can include joint regression layer) can be configured to output joint position based on the posture and shape parameters received at the input. Joint position 214 can be inferred by applying linear regression to one or more mesh vertices, which are determined according to predicted posture and shape parameters 208 and / or SMPL model. Joint position loss 216 can be calculated to represent the difference between the inferred joint position 214 and the input joint position 202. For example, denoting the input joint position 202 as J, and denoting the pose and shape parameters 208 predicted by the ANN 204 as G(J), the inferred joint position 214 may be denoted by X(G(J)), where X may represent the joint position regressor 212. Thus, the inferred joint position 214 may be denoted by The joint position loss 216 is calculated, and the total loss function L for training the ANN 204 can be derived as:
[0028]
[0029] Therefore, ANN 204 can adjust its execution parameters (e.g., weights) with the goal of minimizing the total loss function L. For example, after obtaining the initial pose and shape estimates , the ANN 204 may update its execution parameters via a back-propagation process (e.g., based on the gradient descent of the loss function relative to the current set of parameters). The ANN may then repeat the above prediction and adjustment process until one or more training termination criteria are met (e.g., after completing a predetermined number of training iterations, until the change in the value of the loss function L between consecutive training iterations falls below a predetermined threshold, etc.).
[0030] In an example embodiment, the training of ANN 204 can be performed using annotated 2D joint positions (e.g., using only annotated 2D joint positions or using 2D and 3D joint positions) as input. During training, ANN 204 can predict posture and shape parameters 208 based on the input 2D joint positions, and infer 3D joint positions via joint position regressor 212 based on the predicted posture and shape parameters. In order to verify the accuracy of the predicted parameters 208, ANN 204 can project the inferred 3D joint positions onto the 2D image plane to obtain corresponding 2D joint positions (e.g., 2D coordinates, key points and / or features indicating 2D joint positions). ANN 204 can then compare the projected 2D joint positions with the annotated 2D joint positions received at the input, and adjust the execution parameters of ANN 204 based on a loss function, which is associated with the annotated 2D joint positions and the projected 2D joint positions. For example, the projection can be performed based on a weak perspective camera model, and the 2D joint position can be derived as x=sΠ(RX(β,θ))+t, where can represent global rotations in axis-angle notation, and s may correspond to translation and scaling, respectively, and Π may represent an orthogonal projection. Thus, the training of the ANN 204 may be performed with the goal of minimizing the loss between the annotated input 2D joint positions and the projected 2D joint positions, which may be expressed as Where J represents the annotated input 2D joint positions. As described herein, training can also be supplemented by further considering one or more of parameter loss (e.g., between input pose and / or shape parameters included in the training data and predicted pose and / or shape parameters) or 3D joint position loss (e.g., between inferred 3D joint positions and 3D joint positions included in the training data).
[0031] ANN 204 can also be trained to predict people's posture and shape based on incomplete (e.g., part) knowledge of people's joint positions. As mentioned above, ANN 204 may only have this incomplete or partial knowledge of people's joint positions because some joint positions of people may be blocked, not observed or otherwise unknown to ANN 204. The training of ANN 204 can take this situation into account. For example, during the training of ANN 204, one or more randomly selected subsets of the joint positions included in the training data can be excluded (e.g., marked as unavailable or not observed) from the input of ANN 204, and ANN 204 can be forced to adjust its execution parameters, to adapt to incomplete input (e.g., ANN can be forced to approximate the prediction of given real data, although only with partial information about joint positions).
[0032] Figure 3 is used to train the neural network system described herein (e.g., Figure 1 The pose and shape regressor 104 in Figure 2 204 in . The process 300 may start at 302, and at 304, the neural network system may initialize its execution parameters, such as weights associated with one or more hidden layers (e.g., fully connected layers) of the neural network system. The parameters may be initialized, for example, by sampling one or more probability distributions or based on parameter values of another neural network with a similar architecture. At 306, the neural network system may receive joint position information associated with a human body at an input, and process the joint position information using the initial execution parameters. The neural network system may predict the posture and shape of a human body based on the input joint position information. At 308, the neural network system may determine adjustments to its execution parameters based on a loss function and a gradient descent (e.g., stochastic gradient descent) associated with the loss function (e.g., for minimizing the loss function). The loss function can be implemented based on the mean square error (MSE) or Euclidean distance between the predicted pose and shape parameters and the true data parameters (e.g., which can be paired with the joint position information in the training dataset) and / or the MSE or Euclidean distance between the input joint positions and / or the joint positions inferred from the predicted pose and shape parameters (e.g., as described herein). The loss function can also take into account the L1 norm, the L2 norm, or the L2 norm and the L1 norm.
[0033] At 310, the neural network system may perform adjustments to its current execution parameters, such as via a back-propagation process. At 312, the neural network system may determine whether one or more training termination criteria are met. For example, if the system has completed a predetermined number of training iterations, if the difference between the predicted parameters and the true data parameters is below a predetermined threshold, or if the change in the value of the loss function between two training iterations is below a predetermined threshold, the system may determine that the training termination criteria are met. If it is determined at 312 that the training termination criteria are not met, the system may return to 306. If it is determined at 312 that the training termination criteria are met, the system may end the training process 300 at 314.
[0034] The neural network systems described herein (e.g., Figure 1 The pose and shape regressor 104 in Figure 2 The ANN 204 in the example may be implemented using one or more processors, one or more storage devices, and / or other suitable auxiliary devices (such as display devices, communication devices, input / output devices, etc.). Figure 4400 , which may be a central processing unit (CPU), a graphics processing unit (GPU), a microcontroller, a reduced instruction set computer (RISC) processor, an application specific integrated circuit (ASIC), an application specific instruction set processor (ASIP), a physical processing unit (PPU), a digital signal processor (DSP), a field programmable gate array (FPGA), or any other circuit or processor capable of performing the functions described herein. Neural network system 400 may also include communication circuitry 404, a memory 406, a mass storage device 408, an input device 410, and / or a communication link 412 (e.g., a communication bus), Figure 4 One or more components shown can exchange information through the communication link. The communication circuit 404 can be configured to send and receive information using one or more communication protocols (e.g., TCP / IP) and one or more communication networks, including local area networks (LANs), wide area networks (WANs), the Internet, wireless data networks (e.g., Wi-Fi, 3G, 4G / LTE, or 5G networks). The memory 406 may include a storage medium configured to store machine-readable instructions, which, when implemented, causes the processor 402 to perform one or more functions described herein. Examples of machine-readable media may include volatile or non-volatile memory, including but not limited to semiconductor memory (e.g., electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM)), flash memory, etc.). The mass storage device 408 may include one or more disks, such as one or more built-in hard disks, one or more removable disks, one or more magneto-optical disks, one or more CD-ROM or DVD-ROM disks, etc., on which instructions and / or data may be stored to facilitate the operation of the processor 402. Input device 410 may include a keyboard, a mouse, a voice control input device, a touch-sensitive input device (e.g., a touch screen), etc., for receiving user input to neural network system 400.
[0035] It should be noted that neural network system 400 can operate as a standalone device or can be connected (e.g., networked or clustered) with other computing devices to perform the functions described herein. Figure 4Only one example of each component is shown in the figure, and those skilled in the art will also understand that the neural network system 400 may include multiple instances of one or more of the components shown in the figure. In addition, although the examples are described herein with reference to various types of neural networks, various types of layers, and / or various tasks performed by certain types of neural networks or layers, these references are made for illustrative purposes only and are not intended to limit the scope of the present disclosure. In addition, the operations of the example neural network system are depicted and described herein in a particular order. However, it should be understood that these operations can occur in various orders, simultaneously, and / or with other operations not presented or described herein. And not all operations that the neural network system can perform are depicted and described herein, and not all illustrated operations need to be performed by the system.
[0036] For simplicity of illustration, the operations of the neural network system may have been depicted and described in a particular order. However, it should be understood that these operations can occur in various orders, simultaneously, and / or with other operations not presented or described herein. In addition, it should be noted that not all operations that the neural network system is capable of performing are depicted and described herein. It should also be noted that not all illustrated operations need to be performed by the neural network system.
[0037] Although the present disclosure has been described according to certain embodiments and generally associated methods, the changes and transformations of the embodiments and methods will be apparent to those skilled in the art. Therefore, the above description of exemplary embodiments does not limit the present disclosure. Other changes, substitutions and variations are also possible without departing from the spirit and scope of the present disclosure. In addition, unless otherwise specifically stated, discussions using terms such as "analyze," "determine," "enable," "identify," "modify," etc. refer to the actions and processes of a computer system or similar electronic computing device, which manipulate and transform data represented as physical (e.g., electronic) quantities within registers and memories of a computer system into other data represented as physical quantities within a computer system memory or other such information storage, transmission or display device.
[0038] It should be understood that the above description is intended to be illustrative, rather than restrictive. After reading and understanding the above description, many other embodiments will be apparent to those skilled in the art. Therefore, the scope of the present disclosure should be determined with reference to the appended claims and the full scope of equivalents to which such claims are assigned.
Claims
1. A system for human body posture and shape estimation, comprising one or more processors and one or more storage devices storing instructions, which, when executed by the one or more processors, cause the one or more processors to: Estimation using artificial neural networks; providing the artificial neural network with information related to the position of one or more joints of the person; determining, via the artificial neural network, a first plurality of parameters associated with a posture of the person based on the information regarding the one or more joint positions of the person; determining, via the artificial neural network, a second plurality of parameters associated with the shape of the person based on the information regarding the one or more joint positions of the person; and generating a human body model representing the posture and shape of the person based on the first plurality of parameters and the second plurality of parameters; in, The artificial neural network is trained to determine the first plurality of parameters and the second plurality of parameters, and the training of the artificial neural network comprises: The artificial neural network receives training data, wherein the training data includes joint positions of a human body; The artificial neural network predicts posture and shape parameters associated with the human body based on the joint positions included in the training data; The artificial neural network infers joint positions of the human body based on the predicted posture and shape parameters; and The artificial neural network adjusts one or more of its execution parameters based on a difference between the inferred joint positions and the joint positions included in the training data; wherein the one or more joint positions of the person are a subset of all joint positions of the person; wherein the training data for training the artificial neural network comprises a plurality of joint positions of the human body, and wherein, during the training of the artificial neural network, a subset of the plurality of joint positions is randomly excluded from an input of the artificial neural network, and the artificial neural network is configured to adapt to incomplete inputs by predicting the posture and shape parameters associated with the human body without the excluded subset of the plurality of joint positions.
2. The system according to claim 1, wherein: The training data also includes posture and shape parameters associated with the joint positions of the human body, and wherein the training of the artificial neural network also includes: the artificial neural network adjusting one or more of its execution parameters based on the difference between the predicted posture and shape parameters and the posture and shape parameters included in the training data.
3. The system according to claim 2, wherein: The training data is derived from a parametric human body model associated with the human body.
4. The system according to claim 1, wherein: The information regarding the one or more joint positions of the person indicates three-dimensional (3D) joint positions of the person.
5. The system according to claim 1, wherein: The information regarding the one or more joint positions of the person indicates two-dimensional (2D) joint positions of the person, the training data includes the 2D joint positions of the person, and the training of the artificial neural network includes: The artificial neural network predicts the pose and shape parameters associated with the human body based on the 2D joint positions included in the training data; The artificial neural network uses the predicted pose and shape parameters to infer three-dimensional (3D) joint positions of the human body; The artificial neural network projects the 3D joint positions into image space to obtain corresponding 2D joint positions; and The artificial neural network adjusts the one or more of its execution parameters based on differences between the projected 2D joint positions and the 2D joint positions included in the training data.
6. The system according to claim 5, wherein: The information relating to the one or more joint positions of the person indicates 2D and 3D joint positions of the person, the training data comprises the 2D and 3D joint positions of the person, and the training of the artificial neural network comprises: The artificial neural network predicts the pose and shape parameters associated with the human body based on the 2D and 3D joint positions included in the training data; The artificial neural network uses the predicted pose and shape parameters to infer 3D joint positions of the human body; The artificial neural network projects the 3D joint positions into image space to obtain corresponding 2D joint positions; and The artificial neural network adjusts the one or more of its execution parameters based on differences between the projected 2D joint positions and the 2D joint positions included in the training data and differences between the inferred 3D joint positions and the 3D joint positions included in the training data.
7. The system according to claim 1, wherein: The first plurality of parameters associated with the posture of the person are determined independently of the second plurality of parameters associated with the shape of the person.
8. A method for estimating a posture and shape associated with a person using an artificial neural network, the method comprising: providing the artificial neural network with information relating to the position of one or more joints of the person; determining, via the artificial neural network, a first plurality of parameters associated with a posture of the person based on the information regarding the one or more joint positions of the person; determining, via the artificial neural network, a second plurality of parameters associated with a shape of the person based on the information regarding the one or more joint positions of the person; as well as generating a human body model representing the posture and shape of the person based on the first plurality of parameters and the second plurality of parameters; Wherein, the artificial neural network is trained to determine the first plurality of parameters and the second plurality of parameters, and the training of the artificial neural network comprises: The artificial neural network receives training data, wherein the training data includes joint positions of a human body; The artificial neural network predicts posture and shape parameters associated with the human body based on the joint positions included in the training data; The artificial neural network infers joint positions of the human body based on the predicted posture and shape parameters; and The artificial neural network adjusts one or more of its execution parameters based on a difference between the inferred joint positions and the joint positions included in the training data; wherein the one or more joint positions of the person constitute a subset of all joint positions of the person; wherein the training data for training the artificial neural network comprises a plurality of joint positions of the human body, and wherein, during the training of the artificial neural network, a subset of the plurality of joint positions is randomly excluded from an input of the artificial neural network, and the artificial neural network is configured to adapt to incomplete inputs by predicting the posture and shape parameters associated with the human body without the excluded subset of the plurality of joint positions.
Citation Information
Patent Citations
Systems and methods for human mesh recovery
US11257586B2
Training method of SMPL parameter prediction model, server and storage medium
CN109859296A