Regression-based parametric model-free 3D human body posture and shape prediction method and device

Through a regression-based parametric non-model method, the body part perception module and the Transformer decoding module are used to calculate the posture parameters, and combined with the inverse kinematics principle and the shape regression module, the high computational cost and local minimum problems of the non-model method in parameterized output are solved, and efficient and real-time 3D human posture and shape prediction are achieved, improving the adaptability and accuracy of the model.

CN118710720BActive Publication Date: 2025-10-14ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410856078.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-28
Publication Date
2025-10-14
Estimated Expiration
2044-06-28

AI Technical Summary

Technical Problem

Existing model-free 3D human pose and shape prediction methods have high computational cost and are prone to falling into local minima when parameterizing output, and are difficult to directly handle application scenarios that require parameter input.

Method used

A regression-based parametric non-model method is adopted. By obtaining human image features and weak perspective camera parameters, the body part perception module and Transformer decoding module are used to calculate the absolute rotation and translation information, and the posture parameters are calculated by combining the inverse kinematics principle. The shape parameters are regressed from the posture parameters and 3D vertex information through the shape regression module, and the model is optimized using the L2 loss function.

Benefits of technology

It achieves efficient and real-time parametric non-model 3D human posture and shape prediction, improves the model's ability to capture and reproduce complex human movements, reduces computational costs, and enhances the model's adaptability and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118710720B_ABST
    Figure CN118710720B_ABST
Patent Text Reader

Abstract

The application discloses a regression-based parameterized non-model 3D human body posture and shape prediction method and device. The method comprises the following steps: processing the scaled human body image by using a non-model method, acquiring 3D vertices, image features and weak perspective camera parameters; projecting the 3D vertices to a 2D plane to acquire 2D vertices; inputting the 2D vertex coordinates and the image features into a body part perception sampling module to acquire part perception image features and initial T posture 3D vertex coordinates; connecting the part perception image features and the initial T posture 3D vertex coordinates by using a ConCat operation and inputting the connected part perception image features and the initial T posture 3D vertex coordinates into a body part decoding module to decode and regress to calculate absolute rotation and translation information of each part of the body; calculating relative rotation of each joint relative to a parent joint according to the absolute rotation information of each part of the body to obtain posture parameters; inputting the posture parameters and the 3D vertex coordinates into a shape regression module, converting the 3D vertex coordinates into standard T posture 3D vertex coordinates and inputting the standard T posture 3D vertex coordinates into a series of full connection layers to regress and calculate shape parameters.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer vision and pattern recognition, and particularly relates to a parameterized non-model 3D human pose and shape prediction method and device based on regression. BACKGROUND

[0002] Estimating 3D human pose and shape (HPS) from a single RGB image is a core challenge in computer vision and has wide applications in robotics, computer graphics, and vision. The mainstream learning-based HPS methods can be roughly divided into two categories: model-based methods and non-model methods. Model-based methods represent 3D human mesh by regressing body model parameters, such as SMPL (Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J Black. 2015. SMPL: A skinned multi-person linear model. ACM Trans. Graph. 34, 6 (2015), 1-16.), which utilizes 3D joint angles and shape vectors. Non-model methods directly regress 3D human representations, such as 3D vertex coordinates, implicit surfaces, and voxel grids, from 2D images.

[0003] Non-model based methods have several advantages over model based methods. First, the output vertex coordinates can closely correspond to the original input image, thus the reconstructed mesh of body shape is usually more accurate than model based methods. Second, the regressed vertex / landmark coordinates have the potential to be registered across different parametric models, such as SMPL and GHUM (Hongyi Xu, Eduard Gabriel Bazavan, Andrei Zanfir, William T Freeman, Rahul Sukthankar, and Cristian Sminchisescu. 2020. Ghum&ghuml: Generative 3D human shape and articulated pose models. In IEEE Conf. Comput. Vis. Pattern Recog. 6184-6193.). However, there are many applications that require parametric input, such as character animation (Zhongjin Luo, Shengcai Cai, Jinguo Dong, Ruibo Ming, Liangdong Qiu, Xiaohang Zhan, and Xiaoguang Han. 2023. RaBit: Parametric Modeling of 3D Biped Cartoon Characters with a Topological-consistent Dataset. In IEEE Conf. Comput. Vis. Pattern Recog.) and perceptual shape motion retargeting (Thiago L. Gomes, Renato Martins, Ferreira, Rafael Azevedo, Guilherme Torres, and Erickson R. Nascimento. 2021. A Shape-Aware Retargeting Approach to Transfer Human Motion and Appearance in Monocular Videos. Int. J. Comput. Vis. (29 Apr 2021).). The output of non-model based methods cannot directly handle these cases, thus it is necessary and valuable to efficiently parameterize the output vertex coordinates of non-model based methods.

[0004] Recently, some works have proposed to parameterize non-model-based methods. SMPLify (Federica Bogo, Angjoo Kanazawa, Christoph Lassner, Peter Gehler, Javier Romero, and Michael J Black. 2016. Keep it SMPL: Automatic estimation of 3D human pose and shape from a single image. In Eur. Conf. Comput. Vis. Springer, 561-578.) is a pioneering method that obtains 3D pose and shape parameters by minimizing an objective function that penalizes the error between the 3D model joint projections and detected 2D joints. EFT (Hanbyul Joo, Natalia Neverova, and Andrea Vedaldi. 2021. Exemplar fine-tuning for 3d human model fitting towards in-the-wild 3d human pose estimation. In Int. Conf. 3D. Vis. IEEE, 42-52.) further utilizes a network to overfit 2D observation data. Although these methods are straightforward, they are computationally expensive and prone to local minima. Other works (Junhyeong Cho, Kim Youwang, and Tae-Hyun Oh. 2022. Cross-Attention of Disentangled Modalities for 3D Human Mesh Recovery with Transformers. In Eur. Conf. Comput. Vis. Springer, 342-359.) directly learn parameter representations by a series of fully connected layers that map 3D vertex coordinates to SMPL parameters. However, their performance significantly drops compared to non-model-based methods. VisDB (Chun-Han Yao, Jimei Yang, Duygu Ceylan, Yi Zhou, Yang Zhou, and Ming-Hsuan Yang. 2022. Learning visibility for robust dense human body estimation. In Eur. Conf. Comput. Vis.) leverages the advantages of both optimization-based and regression-based parameterization processes, i.e., regressing SMPL model parameters from all vertex coordinates and then fitting only the visible part of the body.However, this is still not the best solution, and it takes about 100 iterations to converge. SUMMARY

[0005] In order to solve the above technical problems and the deficiencies in the field, the present application provides a regression-based parameterized non-model 3D human pose and shape prediction method, which is a new paradigm based on regression and can parameterize any existing non-model-based 3D human pose and shape prediction method.

[0006] A regression-based parameterized non-model 3D human pose and shape prediction method is a pure regression algorithm that can generate parameterized output from any existing non-model network in real time, thereby bypassing the shortcomings of previous optimization-based processes.

[0007] A regression-based parameterized non-model 3D human pose and shape prediction method includes the following steps:

[0008] S1: Obtain an original human image and scale it to a predetermined size;

[0009] S2: Process the scaled human image using a non-model method to obtain human 3D vertices, image features, and weak perspective camera parameters; project the human 3D vertices onto a 2D plane using the camera parameters to obtain human 2D vertices in the image;

[0010] S3: Input the 2D vertex coordinates and image features into a body part perception sampling module, which divides the human body into 24 parts, each part having its corresponding part 2D vertex and part image feature, and initializes the T pose 3D vertex coordinates from typical SMPL parameters, divides the part perception image features from the overall image features according to the part 2D vertices, and connects the part perception image features and the initialized T pose 3D vertex coordinates using the ConCat operation and inputs them into a body part decoding module based on Transformer, to obtain part perception feature embedding, and then in the world coordinate system, use the part perception feature embedding to regress the absolute rotation and translation information of each part of the body;

[0011] S4: According to the absolute rotation information of each part of the body, calculate the relative rotation of each joint relative to its parent joint by inverse kinematics, to obtain the pose parameters of the SMPL model;

[0012] S5: Input the pose parameters of the SMPL model and the human 3D vertex coordinates obtained in step S2 into a shape regression module, perform inverse linear blend skinning to convert the human 3D vertex coordinates obtained in step S2 into standard T pose 3D vertex coordinates, and then input the standard T pose 3D vertex coordinates into a series of fully connected layers to calculate the shape parameters of the SMPL model.

[0013] In step S3, the L2 loss function L can be used as follows: pnp Optimize the parameters of the regression model to supervise the absolute rotation prediction:

[0014]

[0015] Where i represents the joint index, N = 23, π(·) is the projection function, R(θ i ′) represents the absolute rotation of the body part corresponding to joint i, represents the 3D vertex coordinates of joint i, D i represents the translation information of joint i, Represents the 2D vertex coordinates of joint i.

[0016] In step S4, the absolute rotation can be converted into relative rotation through the inverse transformation of the kinematic tree. Specifically, the relative rotation of each joint relative to its parent joint can be calculated as follows:

[0017]

[0018] in, and R(θ0′) represent the relative rotation of the root joint and the absolute rotation of the body part corresponding to the root joint, respectively. j represents the joint index. Represents the relative rotation of joint j relative to its parent joint, R T (θ p ' (j) ) represents R(θ p ' (j) ), p(j) represents the parent joint index of joint j, R(θ p ' (j) ) represents the absolute rotation of the body part corresponding to the parent joint p(j) of joint j, R(θ j ′) represents the absolute rotation of the body part corresponding to joint j.

[0019] In step S5, the shape parameters of the SMPL model can be calculated as follows:

[0020] J(β,θ)=J reg V 3d ,

[0021]

[0022] in:

[0023] J(β,θ) represents the 3D coordinates of the joint points of the original human body posture in the scaled human body image, which is obtained by mapping the 3D vertices of the human body. reg is the joint regressor, V 3dThe 3D vertex coordinates of the human body obtained in step S2;

[0024] Indicates the relative translation of the root joint, the actual value is 0, represents the transpose of R0(θ), R0(θ) represents the rotation matrix of the root joint of the original human body posture in the scaled human body image, which is used to change the original human body posture into T posture, J0(β,θ) represents the 3D coordinates of the root joint of the original human body posture in the scaled human body image, i represents the joint index, Represents the offset of joint i relative to the root joint, Represents R i The transpose of (θ), R i (θ) represents the rotation matrix of joint i of the original human pose in the scaled human image, J i (β,θ) represents the 3D coordinates of the joint point of joint i in the original human posture in the scaled human image, p(i) represents the parent joint index of joint i, J p(i) (β,θ) represents the 3D coordinates of the joint point p(i) of the parent joint i of the original human posture in the scaled human image;

[0025] J0(β) represents the 3D coordinate of the root joint of T posture, J i (β) represents the 3D coordinate of joint i in T pose, Represents the offset of the parent joint p(i) of joint i relative to the root joint;

[0026] β represents the shape parameter and θ represents the attitude parameter.

[0027] In step S5, the loss function L can be calculated based on the following formula: shape Optimize the parameters of the regression calculation model:

[0028] L shape =L vertices +L betas ,

[0029] Among them, L vertices represents the L2 loss of the real human 3D vertex coordinates and the predicted human 3D vertex coordinates, L betas Represents the L2 loss of the regression shape parameter and the true shape parameter.

[0030] The present invention also provides a computer device comprising a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, and when the computer program is run, the processor executes the regression-based parameterized non-model 3D human body posture and shape prediction method.

[0031] The present invention also provides a computer-readable storage medium, which stores a program or instruction. When the program or instruction is executed by a computer device, the computer device executes the regression-based parameterized non-model 3D human body posture and shape prediction method.

[0032] The present invention also provides a computer program product, comprising a computer program, which, when executed by a computer device, enables the computer device to execute the regression-based parameterized non-model 3D human body posture and shape prediction method.

[0033] As a general inventive concept, the present invention also provides a regression-based parameterized non-model 3D human body pose and shape prediction device, comprising:

[0034] An acquisition module, configured to acquire an original human body image and scale it to a predetermined size;

[0035] The non-model 3D human pose prediction module is used to process the human body image scaled by the acquisition module using a non-model method, obtain the human body's 3D vertices, image features and weak perspective camera parameters, and use the camera parameters to project the human body's 3D vertices onto a 2D plane to obtain the human body's 2D vertices in the image;

[0036] The body part perception sampling module receives the 2D vertex coordinates and image features obtained by the non-model 3D human pose prediction module, divides the human body into 24 parts, each with its corresponding partial 2D vertex and partial image features, and initializes the T-pose 3D vertex coordinates from typical SMPL parameters. The partial perception image features are obtained from the overall image features based on the partial 2D vertices. The partial perception image features and the initialized T-pose 3D vertex coordinates are concatenated using the ConCat operation and input into the body part decoding module based on the Transformer implementation;

[0037] The body part decoding module is used to decode the part-aware feature embeddings and then use the part-aware feature embeddings to regress the absolute rotation and translation information of each body part in the world coordinate system;

[0038] The posture regression module is used to calculate the relative rotation of each joint relative to its parent joint based on the absolute rotation information of each body part obtained by the body part decoding module, and obtain the posture parameters of the SMPL model through the inverse kinematics principle;

[0039] a shape regression module, configured to receive the pose parameters of the SMPL model obtained by the pose regression module and the 3D human body vertex coordinates obtained by the non-model 3D human body pose prediction module, perform inverse linear blend skinning calculation to convert the 3D human body vertex coordinates obtained by the non-model 3D human body pose prediction module into standard T-pose 3D vertex coordinates, and then send the standard T-pose 3D vertex coordinates into a series of fully connected layer regression calculation to obtain the shape parameters of the SMPL model.

[0040] The regression-based parameterized non-model 3D human body pose and shape prediction device can execute the regression-based parameterized non-model 3D human body pose and shape prediction method.

[0041] Compared with the prior art, the present application has the following beneficial effects:

[0042] The present application can parameterize the output of any existing non-model 3D human body pose and shape prediction method. The present application proposes a pose regression paradigm that can effectively utilize the information obtained by the existing non-model method and calculate the pose parameters through motion tree conversion. The present application proposes a shape regression paradigm that uses inverse linear blend skinning to regress the shape parameters from the pose parameters and 3D vertex information. BRIEF DESCRIPTION OF DRAWINGS

[0043] Figure 1 A flowchart of a regression-based parameterized non-model 3D human body pose and shape prediction method in the specific embodiment.

[0044] Figure 2 A schematic diagram of a regression-based parameterized non-model 3D human body pose and shape prediction device in the specific embodiment. DETAILED DESCRIPTION

[0045] The present application will be further described below in conjunction with the drawings and specific embodiments. It should be understood that these embodiments are only used to illustrate the present application and not to limit the scope of the present application.

[0046] In combination Figure 1 A regression-based parameterized non-model 3D human body pose and shape prediction method, comprising the steps of:

[0047] S1: Obtain an original human body image and scale it to a predetermined size.

[0048] For example, OpenCV can be used to read a picture with a person, and then scale it to a specified size (such as 512x512 resolution).

[0049] S2: processing the scaled human body image by using a non-model method to obtain human body 3D vertices, image features and weak perspective camera parameters; projecting the human body 3D vertices onto a 2D plane by using the camera parameters to obtain human body 2D vertices in the image.

[0050] The non-model method used in the application can be any existing non-model method. For example, the human body 3D vertices V 3d ∈R 6890×3 , the human body 2D vertices V 2d ∈R 6890×2 . Specifically, the FastMETRO model can be used as a non-model 3D human body pose prediction method, and the picture processed in step S1 is input into the FastMETRO to obtain human body 3D vertices, image features and weak perspective camera parameters. The human body 3D vertices can be projected onto a 2D plane by using the camera parameters to obtain human body 2D vertex coordinate information in the picture.

[0051] S3: inputting the 2D vertex coordinates and image features into a body part perception sampling module, the body part perception sampling module divides the human body into 24 parts, each part has its corresponding part 2D vertex and part image feature, and the T pose 3D vertex coordinates are initialized from the typical SMPL parameters, the part 2D vertex is used to divide the part perception image feature from the overall image feature, the part perception image feature and the initialized T pose 3D vertex coordinates are connected by using the ConCat operation and input into the body part decoding module based on the Transformer to obtain the part perception feature embedding, and then the absolute rotation and translation information of each part of the body are calculated by using the part perception feature embedding in the world coordinate system.

[0052] The SMPL mesh is divided into 24 parts, which correspond to the influence of the 24 rotation vectors of the SMPL model.

[0053] In step S3, the L2 loss function L pnp is used for parameter optimization of the regression calculation model to supervise the absolute rotation prediction:

[0054]

[0055] wherein i represents the joint index, N = 23, π(·) is a projection function (perspective projection when the camera internal parameter is available, or weak perspective projection on an uncalibrated image), R(θ i ′) represents the absolute rotation of the body part corresponding to the joint i, represents the 3D vertex coordinates of the joint i, D i represents the translation information of the joint i, represents the 2D vertex coordinates of the joint i.

[0056] The above method combines precise geometric transformation and the powerful function of deep learning, improving the ability of the model to capture and reproduce complex human body movements.

[0057] S4: According to the absolute rotation information of each part of the body, the relative rotation of each joint relative to its parent joint is calculated by inverse kinematics principle, and the pose parameters of the SMPL model are obtained.

[0058] In step S4, the relative rotation of each joint relative to its parent joint can be calculated as follows:

[0059]

[0060] wherein, and R(θ0') represent the relative rotation of the root joint and the absolute rotation of the root joint corresponding to the body part, j represents the joint index, R(θj') represents the relative rotation of joint j relative to its parent joint, R T (θ p ′ (j) ) represents the transpose of R(θ p ′ (j) ), p(j) represents the parent joint index of joint j, R(θ p ′ (j) ) represents the absolute rotation of the parent joint p(j) of joint j corresponding to the body part, and R(θ j ') represents the absolute rotation of joint j corresponding to the body part.

[0061] S5: The pose parameters of the SMPL model and the 3D vertex coordinates of the human body obtained in step S2 are input into the shape regression module, inverse linear blending skinning calculation is performed to convert the 3D vertex coordinates of the human body obtained in step S2 into standard T-pose 3D vertex coordinates, and then the standard T-pose 3D vertex coordinates are input into a series of fully connected layer regression calculation to obtain the shape parameters of the SMPL model.

[0062] In step S5, the shape parameters of the SMPL model can be calculated as follows:

[0063] J(β,θ)=J reg V 3d ,

[0064]

[0065] wherein:

[0066] J(β,θ) represents the 3D joint coordinates of the original human pose in the scaled human image, which is obtained by mapping the 3D vertex, J reg is the joint regressor, V 3d is the 3D vertex coordinates of the human body obtained in step S2;

[0067] Indicates the relative translation of the root joint, the actual value is 0, represents the transpose of R0(θ), R0(θ) represents the rotation matrix of the root joint of the original human body posture in the scaled human body image, which is used to change the original human body posture into T posture, J0(β,θ) represents the 3D coordinates of the root joint of the original human body posture in the scaled human body image, i represents the joint index, Represents the offset of joint i relative to the root joint, Represents R i The transpose of (θ), R i (θ) represents the rotation matrix of joint i of the original human pose in the scaled human image, J i (β,θ) represents the 3D coordinates of the joint point of joint i in the original human posture in the scaled human image, p(i) represents the parent joint index of joint i, J p(i) (β,θ) represents the 3D coordinates of the joint point p(i) of the parent joint i of the original human posture in the scaled human image;

[0068] J0(β) represents the 3D coordinate of the root joint of T posture, J i (β) represents the 3D coordinate of joint i in T pose, Represents the offset of the parent joint p(i) of joint i relative to the root joint;

[0069] β represents the shape parameter and θ represents the attitude parameter.

[0070] In step S5, the loss function L can be calculated based on the following formula: shape Optimize the parameters of the regression calculation model:

[0071] L shape =L vertices +L betas ,

[0072] Among them, L vertices represents the L2 loss of the real human 3D vertex coordinates and the predicted human 3D vertex coordinates, L betas Represents the L2 loss of the regression shape parameter and the true shape parameter.

[0073] The above training strategy not only helps to improve the accuracy of the shape, but also enhances the model's adaptability to complex postures and deformations, making it better suitable for practical application scenarios.

[0074] A computer device includes a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory. When the computer program is run, the processor executes the regression-based parameterized non-model 3D human body posture and shape prediction method.

[0075] A computer-readable storage medium stores a program or instruction, which, when executed by a computer device, causes the computer device to execute the regression-based parameterized non-model 3D human body pose and shape prediction method.

[0076] A computer program product includes a computer program. When the computer program is executed by a computer device, the computer device is caused to perform the regression-based parameterized non-model 3D human body pose and shape prediction method.

[0077] Combine Figure 2 , a regression-based parameterized model-free 3D human pose and shape prediction device, comprising:

[0078] An acquisition module 21 is used to acquire an original human body image and scale it to a predetermined size;

[0079] A non-model 3D human posture prediction module 22 is used to process the human body image scaled by the acquisition module 21 using a non-model method, obtain the human body's 3D vertices, image features, and weak perspective camera parameters, and use the camera parameters to project the human body's 3D vertices onto a 2D plane to obtain the human body's 2D vertices in the image;

[0080] The body part perception sampling module 23 is used to receive the 2D vertex coordinates and image features obtained by the non-model 3D human pose prediction module 22, divide the human body into 24 parts, each part has its corresponding partial 2D vertex and partial image features, and initialize the T-pose 3D vertex coordinates from typical SMPL parameters. The partial perception image features are obtained from the overall image features based on the partial 2D vertices. The partial perception image features and the initialized T-pose 3D vertex coordinates are concatenated using the ConCat operation and input into the body part decoding module 24 based on the Transformer implementation;

[0081] The body part decoding module 24 is used to decode and obtain part-aware feature embeddings, and then use the part-aware feature embeddings to perform regression calculations on the absolute rotation and translation information of each body part in the world coordinate system;

[0082] The posture regression module 25 is used to calculate the relative rotation of each joint relative to its parent joint based on the absolute rotation information of each body part obtained by the body part decoding module 24, and obtain the posture parameters of the SMPL model through the inverse kinematics principle;

[0083] The shape regression module 26 is used to receive the posture parameters of the SMPL model obtained by the posture regression module 25 and the 3D vertex coordinates of the human body obtained by the non-model 3D human posture prediction module 22, perform inverse linear mixed skinning calculation to convert the 3D vertex coordinates of the human body obtained by the non-model 3D human posture prediction module 22 into standard T posture 3D vertex coordinates, and then send the standard T posture 3D vertex coordinates to a series of fully connected layer regression calculations to obtain the shape parameters of the SMPL model.

[0084] The above-mentioned regression-based parameterized non-model 3D human body pose and shape prediction device can execute the above-mentioned regression-based parameterized non-model 3D human body pose and shape prediction method when running.

[0085] In addition, it should be understood that after reading the above description of the present invention, those skilled in the art may make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the claims attached to this application.

Claims

1. A regression-based parametric model-free 3D human pose and shape prediction method, characterized in that: Including steps: S1: Obtain the original human body image and scale it to a predetermined size; S2: Use a non-model method to process the scaled human body image to obtain the human body's 3D vertices, image features, and weak perspective camera parameters; use the camera parameters to project the human body's 3D vertices onto a 2D plane to obtain the human body's 2D vertices in the image; S3: Input the 2D vertex coordinates and image features into the body part-aware sampling module. The body part-aware sampling module divides the human body into 24 parts, each of which has its corresponding partial 2D vertex and partial image features. The T-pose 3D vertex coordinates are initialized from typical SMPL parameters. Partially perceived image features are obtained from the overall image features based on the partial 2D vertices. The partially perceived image features and the initialized T-pose 3D vertex coordinates are concatenated using the ConCat operation and input into the body part decoding module based on the Transformer implementation. The decoding obtains the part-aware feature embedding. Then, the part-aware feature embedding is used to regress the absolute rotation and translation information of each body part in the world coordinate system. S4: Based on the absolute rotation information of each body part, the relative rotation of each joint relative to its parent joint is calculated through the inverse kinematics principle to obtain the posture parameters of the SMPL model; S5: The posture parameters of the SMPL model and the 3D vertex coordinates of the human body obtained in step S2 are input into the shape regression module, and the inverse linear mixed skinning calculation is performed to convert the 3D vertex coordinates of the human body obtained in step S2 into the standard T posture 3D vertex coordinates. The standard T posture 3D vertex coordinates are then sent to a series of fully connected layer regression calculations to obtain the shape parameters of the SMPL model.

2. The regression-based parameterized non-model 3D human pose and shape prediction method according to claim 1, characterized in that: In step S3, based on the L2 loss function L pnp Optimize the parameters of the regression model to supervise the absolute rotation prediction: Where i represents the joint index, N = 23, π(·) is the projection function, R(θ i ′ ) represents the absolute rotation of the body part corresponding to joint i, represents the 3D vertex coordinates of joint i, D i represents the translation information of joint i, Represents the 2D vertex coordinates of joint i.

3. The regression-based parameterized non-model 3D human pose and shape prediction method according to claim 1, characterized in that: In step S4, the relative rotation of each joint relative to its parent joint is calculated as follows: in, and R(θ0 ′ ) represent the relative rotation of the root joint and the absolute rotation of the body part corresponding to the root joint, j represents the joint index, Represents the relative rotation of joint j relative to its parent joint, R T (θ p ′ (j) ) represents R(θ p ′ (j) ), p(j) represents the parent joint index of joint j, R(θ p ′ (j) ) represents the absolute rotation of the body part corresponding to the parent joint p(j) of joint j, R(θ j ′ ) represents the absolute rotation of the body part corresponding to joint j.

4. The regression-based parameterized non-model 3D human pose and shape prediction method according to claim 1, characterized in that In step S5, the shape parameters of the SMPL model are calculated as follows: J(β,θ)=J reg V 3d , in: J(β,θ) represents the 3D coordinates of the joint points of the original human body posture in the scaled human body image, which is obtained by mapping the 3D vertices of the human body. reg is the joint regressor, V 3d The 3D vertex coordinates of the human body obtained in step S2; Indicates the relative translation of the root joint, the actual value is 0, represents the transpose of R0(θ), R0(θ) represents the rotation matrix of the root joint of the original human body posture in the scaled human body image, which is used to change the original human body posture into T posture, J0(β,θ) represents the 3D coordinates of the root joint of the original human body posture in the scaled human body image, i represents the joint index, Represents the offset of joint i relative to the root joint, Represents R i The transpose of (θ), R i (θ) represents the rotation matrix of joint i of the original human pose in the scaled human image, J i (β,θ) represents the 3D coordinates of the joint point of joint i in the original human posture in the scaled human image, p(i) represents the parent joint index of joint i, J p(i) (β,θ) represents the 3D coordinates of the joint point p(i) of the parent joint i of the original human posture in the scaled human image; J0(β) represents the 3D coordinate of the root joint of T posture, J i (β) represents the 3D coordinate of joint i in T pose, Represents the offset of the parent joint p(i) of joint i relative to the root joint; β represents the shape parameter and θ represents the attitude parameter.

5. The regression-based parameterized model-free 3D human pose and shape prediction method according to claim 1, characterized in that: In step S5, based on the loss function L shape Optimize the parameters of the regression calculation model: L shape =L vertices +L betas , Among them, L vertices represents the L2 loss of the real human 3D vertex coordinates and the predicted human 3D vertex coordinates, L betas Represents the L2 loss of the regression shape parameter and the true shape parameter.

6. A computer device comprising a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, wherein: When the computer program is executed, the processor is enabled to execute the regression-based parameterized non-model 3D human body posture and shape prediction method according to any one of claims 1 to 5.

7. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a program or instruction, and when the program or instruction is executed by a computer device, the computer device executes the regression-based parameterized non-model 3D human body posture and shape prediction method according to any one of claims 1 to 5.

8. A computer program product comprising a computer program, characterized in that When the computer program is executed by a computer device, the computer device is caused to execute the regression-based parameterized non-model 3D human body pose and shape prediction method according to any one of claims 1 to 5.

9. A regression-based parameterized non-model 3D human body posture and shape prediction device, characterized in that: include: An acquisition module, configured to acquire an original human body image and scale it to a predetermined size; The non-model 3D human pose prediction module is used to process the human body image scaled by the acquisition module using a non-model method, obtain the human body's 3D vertices, image features and weak perspective camera parameters, and use the camera parameters to project the human body's 3D vertices onto a 2D plane to obtain the human body's 2D vertices in the image; The body part perception sampling module receives the 2D vertex coordinates and image features obtained by the non-model 3D human pose prediction module, divides the human body into 24 parts, each with its corresponding partial 2D vertex and partial image features, and initializes the T-pose 3D vertex coordinates from typical SMPL parameters. The partial perception image features are obtained from the overall image features based on the partial 2D vertices. The partial perception image features and the initialized T-pose 3D vertex coordinates are concatenated using the ConCat operation and input into the body part decoding module based on the Transformer implementation; The body part decoding module is used to decode the part-aware feature embeddings and then use the part-aware feature embeddings to regress the absolute rotation and translation information of each body part in the world coordinate system; The posture regression module is used to calculate the relative rotation of each joint relative to its parent joint based on the absolute rotation information of each body part obtained by the body part decoding module, and obtain the posture parameters of the SMPL model through the inverse kinematics principle; The shape regression module is used to receive the posture parameters of the SMPL model obtained by the posture regression module and the 3D vertex coordinates of the human body obtained by the non-model 3D human posture prediction module, perform inverse linear mixed skinning calculation to convert the 3D vertex coordinates of the human body obtained by the non-model 3D human posture prediction module into standard T posture 3D vertex coordinates, and then send the standard T posture 3D vertex coordinates to a series of fully connected layer regression calculations to obtain the shape parameters of the SMPL model.

Citation Information

Patent Citations

  • 3D human body posture estimation method based on complementary enhancement of key points and grid vertexes

    CN116631064A

  • Three-dimensional human body posture reconstruction method based on multi-camera self-calibration technology

    CN117635843A