Multi-view three-dimensional human body posture optimization reconstruction method and system based on human body skeleton prior

Through the multi-eye three-dimensional human posture optimization reconstruction method based on the human skeleton prior, the triangulation method and skeleton optimization rules, the accuracy and robustness of three-dimensional human posture reconstruction in the actual environment in the existing technology are solved, and high-precision three-dimensional human posture optimization is achieved.

CN120279096APending Publication Date: 2025-07-08SHANGHAI JIAOTONG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510411477.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing two-dimensional binocular three-dimensional human posture optimization reconstruction method based on deep learning has deteriorated performance in actual environments, especially when the camera position difference is large and the human posture movement is complex, significant estimation errors occur.

Method used

The multi-eye three-dimensional human posture optimization reconstruction method based on the human skeleton prior is adopted, and the triangular method is used to calculate and optimize the three-dimensional human posture reconstruction results through the triangular method, combined with the minimization of multi-eye reprojection error.

Benefits of technology

In a variable practical environment, the robustness and accuracy of three-dimensional human posture reconstruction are improved, and the practical level requirement of the average error of the whole body degree of freedom is less than 5cm, which enhances the applicability of the algorithm in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279096A_ABST
    Figure CN120279096A_ABST
Patent Text Reader

Abstract

The invention provides a multi-view three-dimensional human body posture optimization reconstruction method and system based on human body skeleton prior, and the method comprises the steps: S1, obtaining a multi-view human body image, and obtaining a multi-view two-dimensional human body posture estimation result based on the obtained multi-view human body image; the obtained multi-view two-dimensional human body posture estimation results are grouped in pairs, and a plurality of groups of candidate three-dimensional human body posture reconstruction results under binocular vision are generated through triangulation calculation and introduction of coordinate point offset; s2, on the basis of the candidate three-dimensional human body posture reconstruction results under the multiple groups of binocular vision, calculating three-dimensional posture rationality error scores through a human body skeleton prior optimization rule, and selecting the three-dimensional human body posture with the minimum error score as a three-dimensional human body posture reconstruction optimization result under the multiple groups of binocular vision; and S3, based on the three-dimensional human body posture estimation optimization result under multiple groups of binocular vision, a final three-dimensional human body posture reconstruction result is obtained through optimization by a multi-view re-projection error minimization method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of three-dimensional human pose optimization, and specifically, to a multi-view three-dimensional human pose optimization and reconstruction method and system based on human skeleton prior knowledge. Background Art

[0002] Three-dimensional human pose estimation is an important task in the field of computer vision and has wide applications in fields such as human-computer interaction, intelligent monitoring, intelligent motion analysis, and medical rehabilitation. Three-dimensional human pose estimation is to recover the three-dimensional pose information of the human body from the given input images or videos. Traditional methods are mainly divided into generative algorithms and discriminant algorithms based on template matching. The model-based generative method mainly considers using a model to represent the human body structure and models the spatial relationship between adjacent parts through the constraints between human body structures. The discriminant method regards the pose estimation task as a regression problem, maps the input image to the feature space through a manually designed feature extraction algorithm, and then learns the mapping from the feature space to the three-dimensional human pose space.

[0003] Three-dimensional human pose estimation based on deep learning can be divided into one-stage and two-stage algorithms in terms of algorithm structure. The one-stage algorithm directly estimates the coordinates of three-dimensional joints from the input RGB image to achieve end-to-end three-dimensional pose estimation. The two-stage algorithm usually extracts features from the input image first, predicts the two-dimensional coordinates or other information of the joints in the pixel space of the image, and then further realizes the transformation from the intermediate result to the three-dimensional space, so as to finally estimate the three-dimensional pose.

[0004] Three-dimensional human pose estimation based on deep learning can be divided into monocular and multi-view algorithms in terms of the number of sensors. Compared with the monocular input algorithm, the multi-view input algorithm can make more full use of the geometric information of multi-view images and obtain more accurate results in the three-dimensional human pose estimation.

[0005] Existing binocular three-dimensional human pose optimization and reconstruction methods based on deep learning mainly adopt a data-driven deep learning method. Limited by the fixed camera poses of existing public datasets, when applied to actual environments with large pose differences from the dataset, the performance will significantly decline. At the same time, limited by the limited human poses and action sequences included in existing public datasets, when applied to human pose actions that are more likely to be occluded, confused, and not covered by other public datasets, significant estimation errors are likely to occur.

[0006] Patent Document CN116152439A (Application No.: 202310191078.5) discloses a method and system for 3D human pose reconstruction based on multi-view human images, belonging to the field of computer vision. The method includes: acquiring human images from multiple perspectives through multiple cameras; determining the depth values of each human surface point by using a pre-trained encoding and decoding network according to the human images from multiple perspectives, the minimum depth value, the maximum depth value, the internal parameter matrix and the external parameter matrix of each camera; determining the human point cloud data according to the depth values of each human surface point, the internal parameter matrix and the external parameter matrix of each camera; determining the 3D key point coordinates of the human body in the oriented bounding box coordinate system by using a pre-trained feature extraction network based on the human point cloud data; and converting the 3D key point coordinates of the human body in the oriented bounding box coordinate system to the camera coordinate system to obtain the 3D key point coordinates of the human body in the camera coordinate system, so as to determine the 3D human pose.

[0007] The multi-view 3D human pose optimization and reconstruction method and system based on human skeleton prior proposed by the present invention optimize and reconstruct the significantly incorrect human poses estimated by deep learning methods in the actual scene, so that the present invention meets the practical requirements in the actual scene where the camera placement poses are flexible and changeable: the 3D human pose reconstruction performance with an average error of less than 5 cm in the whole body degree of freedom significantly enhances the robustness of the 3D human pose reconstruction of the present invention in the actual environment with more diverse human pose actions. Summary of the Invention

[0008] Aiming at the defects in the prior art, the purpose of the present invention is to provide a multi-view 3D human pose optimization and reconstruction method and system based on human skeleton prior.

[0009] According to a multi-view 3D human pose optimization and reconstruction method based on human skeleton prior provided by the present invention, it includes:

[0010] Step S1: Obtain multi-view human images, and obtain multi-view 2D human pose estimation results based on the obtained multi-view human images; group the obtained multi-view 2D human pose estimation results in pairs, and generate multiple groups of candidate 3D human pose reconstruction results under binocular vision by calculating through triangulation and introducing coordinate point offsets.

[0011] Step S2: Based on multiple groups of candidate 3D human pose reconstruction results under binocular vision, calculate the 3D pose rationality error score through the human skeleton prior optimization rule, and select the 3D human pose with the smallest error score as the 3D human pose reconstruction optimization result under multiple groups of binocular vision.

[0012] Step S3: Based on the optimized results of the 3D human pose estimation under multiple groups of binocular vision, optimize to obtain the final 3D human pose reconstruction result through the multi-view reprojection error minimization method.

[0013] Preferably, the step S1 includes:

[0014] Step S1.1: Obtain multi-view human body images based on the acquisition method of synchronization and calibration of any multi-view camera;

[0015] Step S1.2: Input the multi-view human body images into a monocular two-dimensional human pose estimation network in sequence to obtain multi-view two-dimensional human poses under the multi-view human body images;

[0016] Step S1.3: Denote the multi-view human body images as C1, C2, C3... Cn, where n is the number of multi-view cameras, and take every two adjacent multi-view cameras as a group to obtain multiple pairs of binocular images {C1, C2}, {C2, C3}... {Cn-1, Cn}. Denote the image with a smaller serial number in each group as the left image and the other image as the right image to form n-1 pairs of left and right images;

[0017] Step S1.4: For the two-dimensional human pose estimation result of the right image in the two-dimensional human poses of each pair of left and right images, perform coordinate point offset along the binocular epipolar line direction that meets the preset requirements to obtain multiple two-dimensional human pose estimation coordinate results of the right image;

[0018] Step S1.5: Based on the fixed two-dimensional human pose estimation result of the left image, calculate multiple three-dimensional human pose estimation results through triangulation respectively with multiple two-dimensional human pose estimation coordinate results of the right image.

[0019] Preferably, the step S2 includes:

[0020] Step S2.1: According to the prior knowledge of human bone length, the estimated results of symmetric human limbs should have the same length. Calculate the symmetric limb length error for the three-dimensional key point coordinates of each three-dimensional human pose estimation result:

[0021]

[0022] where R is the set of limbs with symmetry in the three-dimensional human skeleton, x1, x2 are any pair of symmetric limbs in this set, L is the limb length calculated from the coordinates of the two end bone key points of this limb in the three-dimensional Euclidean coordinate system, and abs represents the absolute value;

[0023] Step S2.2: Based on multiple three-dimensional human pose estimation results, calculate their limb length ratios respectively, and further calculate the absolute difference between the limb length ratio and the reasonable range of the prior knowledge of the limb length ratios of the human body. Sum the differences obtained for all limbs and denote it as the prior error of the limb length ratio;

[0024]

[0025] Among them, G is the set of limbs of the three-dimensional human skeleton except for the upper body trunk of the human body used as a reference, x is any limb in this set, y is the limb of the human body trunk used as the reference for L, L is the limb length calculated from the coordinates of the two key bone points at both ends of each limb in the three-dimensional Euclidean coordinate system, P represents the mean value of the prior of the human limb length ratio of the x limb, abs represents the absolute value, and k is the weight hyperparameter;

[0026] Step S2.3: For multiple three-dimensional human pose estimation results, calculate the radian values of their joint motion angles respectively, and take the absolute difference with the radian values of the prior range of human kinematics. Sum the differences obtained for all limbs, and denote it as the kinematic prior error;

[0027]

[0028] Among them, r is the radian value of the angle calculated in the three-dimensional Euclidean space coordinate system of two limbs connected by any three-dimensional human bone key point, V is the range of the human motion angle corresponding to this human bone key point in the prior of the human motion angle, V min and V max represent the two radian edge values, one large and one small, of V, and min() is the operation of taking the smaller value;

[0029] Step S2.4: Calculate the three-dimensional pose rationality error score:

[0030] Loss = α × Loss1 + b × Loss2 + c × Loss3

[0031] Among them, a, b, and c are the weighted hyperparameters of the scores of the three parts;

[0032] Step S2.5: Calculate the three-dimensional pose rationality error scores for multiple candidate three-dimensional human poses respectively, and finally select the three-dimensional human pose corresponding to the smallest three-dimensional pose rationality error score as the optimized reconstruction result of the three-dimensional human pose of the left and right image pairs.

[0033] Preferably, the reasonable range of the prior of the length ratio of each human limb includes: Denote the change range of the ratio of each human limb to the length of the human body trunk used as the reference limb as the prior range of the human limb length ratio;

[0034] Based on the true values of the three-dimensional human poses of multiple collectors of different genders, ages, and body shapes in multiple public datasets and data collected in actual scenarios, the initial values are statistically obtained. The initial values include the mean and variance of the ratio of the length of each human limb to the trunk length, and the reasonable range of the prior of the length ratio of each human limb is constructed based on the Gaussian distribution.

[0035] Preferably, the radian values of the prior range of human kinematics include: According to the prior knowledge of the human skeleton, each limb of the human body has biologically restricted movement angles on the human skeleton, denoted as the prior of human movement angles; the prior range of human movement angles is statistically obtained based on the three-dimensional human body pose ground truth of multiple collectors of different genders, ages, and body sizes in multiple public datasets and actual scenario collected data, including the movement boundaries and reachable areas of human joint angles.

[0036] Preferably, step S3 includes: Based on the optimization results of three-dimensional human body pose estimation under multiple binocular visions, calculate the reprojection error of each group of optimization results of three-dimensional human body pose estimation under binocular vision in other groups of binocular visions, and obtain the final optimized reconstruction result of multi-view three-dimensional human body pose through the softmax weighted average method.

[0037] A multi-view three-dimensional human body pose optimization and reconstruction method based on prior knowledge of the human skeleton according to the present invention includes:

[0038] Module M1: Obtain multi-view human body images, and obtain multi-view two-dimensional human body pose estimation results based on the obtained multi-view human body images; group the obtained multi-view two-dimensional human body pose estimation results in pairs, and generate multiple groups of candidate three-dimensional human body pose reconstruction results under binocular vision through triangulation calculation and introduction of coordinate point offset;

[0039] Module M2: Based on multiple groups of candidate three-dimensional human body pose reconstruction results under binocular vision, calculate the three-dimensional pose rationality error score through the prior optimization rule of the human skeleton, and select the three-dimensional human body pose with the smallest error score as the optimized reconstruction result of the three-dimensional human body pose under multiple groups of binocular vision;

[0040] Module M3: Based on the optimization results of three-dimensional human body pose estimation under multiple groups of binocular vision, optimize to obtain the final three-dimensional human body pose reconstruction result through the multi-view reprojection error minimization method.

[0041] Preferably, module M1 includes:

[0042] Module M1.1: Obtain multi-view human body images through the synchronous and calibrated acquisition method of any multi-view camera;

[0043] Module M1.2: Input the multi-view human body images into the single-view two-dimensional human body pose estimation network in sequence to obtain the multi-view two-dimensional human body poses under the multi-view human body images;

[0044] Module M1.3: Denote the multi-view human body images as C1, C2, C3... Cn, where n is the number of multi-view cameras, and take every two adjacent multi-view cameras as a group to obtain multiple pairs of binocular images {C1, C2}, {C2, C3}... {Cn-1, Cn}. Denote the image with a smaller serial number in each group as the left image and the other image as the right image to form n - 1 pairs of left and right images;

[0045] Module M1.4: For the estimated result of the two-dimensional human body pose of the right image in each pair of left and right images, perform a coordinate point offset along the binocular epipolar line direction that meets the preset requirements to obtain multiple estimated coordinate results of the two-dimensional human body pose of the right image;

[0046] Module M1.5: Based on the fixed estimated result of the two-dimensional human body pose of the left image, calculate multiple estimated results of the three-dimensional human body pose respectively through triangulation with multiple estimated coordinate results of the two-dimensional human body pose of the right image.

[0047] Preferably, the module M2 includes:

[0048] Module M2.1: According to the prior knowledge of human bone length, the estimated results of symmetric human limbs should have the same length. Calculate the symmetric limb length error for the three-dimensional key point coordinates of each estimated result of the three-dimensional human body pose:

[0049]

[0050] Among them, R is the set of limbs where the three-dimensional human skeleton has symmetry, x1, x2 are any pair of symmetric limbs in this set, L is the limb length calculated from the coordinates of the two end bone key points of this limb in the three-dimensional Euclidean coordinate system, and abs represents the absolute value;

[0051] Module M2.2: Based on multiple estimated results of the three-dimensional human body pose, calculate their limb length ratios respectively, and further calculate the absolute difference between the limb length ratio and the reasonable range of the prior knowledge of the human limb length ratio. Sum the differences obtained for all limbs, which is denoted as the prior error of the limb length ratio;

[0052]

[0053] Among them, G is the set of limbs of the three-dimensional human skeleton except for the upper body torso of the human body as the benchmark, x is any limb in this set, y is the limb of the human torso as the benchmark of L, L is the limb length calculated from the coordinates of the two end bone key points of each limb in the three-dimensional Euclidean coordinate system, P represents the mean value of the prior knowledge of the human limb length ratio of the x limb, abs represents the absolute value, and k is the weight hyperparameter;

[0054] Module M2.3: For multiple three-dimensional human pose estimation results, calculate the radian values of the joint movement angles respectively, take the absolute difference with the radian values within the prior range of human kinematics, and sum up the differences for all limbs, which is denoted as the prior kinematic error;

[0055]

[0056] Among them, r is the radian value of the angle calculated in the three-dimensional Euclidean space coordinate system for two limbs connected by any three-dimensional human bone key point, V is the range of human movement angles corresponding to this human bone key point in the prior human movement angles, V min and V max represent the two radian edge values of V, one larger and one smaller, and min() is the operation of taking the smaller value;

[0057] Module M2.4: Calculate the rationality error score of the three-dimensional pose:

[0058] Loss = α × Loss1 + b × Loss2 + c × Loss3

[0059] Among them, a, b, and c are the weighted hyperparameters for the scores of the three parts;

[0060] Module M2.5: Calculate the rationality error scores of the three-dimensional poses for multiple candidate three-dimensional human poses respectively, and finally select the three-dimensional human pose corresponding to the smallest rationality error score of the three-dimensional pose as the optimized reconstruction result of the three-dimensional human pose for the left and right image pairs.

[0061] Preferably, the module M3 includes: Based on the optimized results of three-dimensional human pose estimation under multiple groups of binocular vision, calculate the reprojection error of each group of optimized results of three-dimensional human pose estimation under binocular vision in other groups of binocular vision respectively, and calculate the final optimized reconstruction result of the multi-view three-dimensional human pose through the softmax weighted average method.

[0062] Compared with the prior art, the present invention has the following beneficial effects:

[0063] 1. In the actual application scenario, due to the significantly incorrect pose reconstruction caused by the variable human clothing, actual actions and poses, the human bone prior is used for optimization, which enhances the three-dimensional human pose reconstruction performance of the multi-view three-dimensional human pose reconstruction method under various camera placement poses in the actual environment, improves the robustness of the multi-view three-dimensional human pose reconstruction method in the actual environment, and meets the practical requirement that the average error of the full-body degrees of freedom is less than 5 cm;

[0064] 2. The multi-view three-dimensional human pose optimization and reconstruction method based on the prior of human skeleton length ratio of the present invention uses the human skeleton length ratio to optimize the abnormal limb length results of multi-view three-dimensional human pose optimization and reconstruction in the actual scene; the multi-view three-dimensional human pose optimization and reconstruction method based on the prior of human symmetry uses the prior of human symmetry to optimize the incorrect results of human symmetry imbalance in multi-view three-dimensional human pose optimization and reconstruction in the actual scene; the multi-view three-dimensional human pose optimization and reconstruction method based on the prior of human kinematics uses the prior of human motion angle to optimize the significantly incorrect results of multi-view three-dimensional human pose optimization and reconstruction in the actual scene.

[0065] 3. The human skeleton prior optimization method of the present invention is not limited to the pose placement of multi-view cameras in the actual scene, and has good three-dimensional human pose optimization and reconstruction effects in various usage environments such as short baseline and long baseline; it is not limited by the limited data of deep learning based on public datasets of mainstream multi-view three-dimensional human pose reconstruction methods, and significantly enhances the robustness of the algorithm in complex scenes with various human poses and limb movements.

[0066] 4. The multi-view three-dimensional human pose reconstruction network based on deep learning of the present invention can be adjusted according to the actual application scenario and resource constraints, and the human skeleton prior optimization method can be adjusted according to the model of human pose key points and the pose representation in three-dimensional space, and has strong generalization. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Other features, objects, and advantages of the present invention will become more apparent by reading the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0068] Figure 1 It is a flowchart of a multi-view three-dimensional human pose optimization and reconstruction method based on human skeleton prior. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0069] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several changes and improvements without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0070] Embodiment 1

[0071] According to a multi-view three-dimensional human pose optimization and reconstruction method based on human skeleton prior provided by the present invention, as Figure 1 shown, it includes:

[0072] Step 1: Based on the multi-view human body images collected by the multi-view vision system, obtain the multi-view two-dimensional human body pose estimation results through any two-dimensional human body pose estimation network. Subsequently, group the multi-view results in pairs, and generate multiple groups of candidate three-dimensional human body pose reconstruction results under binocular vision through triangulation calculation and introducing coordinate point offsets.

[0073] Step 2: Based on multiple groups of candidate three-dimensional human body pose reconstruction results, calculate the three-dimensional pose rationality error score through the human bone prior optimization rule, and select the three-dimensional human body pose with the smallest error score as the optimized result of the three-dimensional human body pose reconstruction under multiple groups of binocular vision.

[0074] Step 3: Based on the optimized results of the three-dimensional human body pose estimation under multiple groups of binocular vision, optimize to obtain the final three-dimensional human body pose reconstruction result through the multi-view reprojection error minimization method.

[0075] Specifically, the Step 1 includes:

[0076] Step 1.1: Input of multi-view human body images. In this embodiment, the required multi-view human body images can use the multi-view image files collected by any multi-view camera after synchronization and calibration as input;

[0077] Step 1.2: Multi-view two-dimensional human body pose estimation. Input the multi-view human body images into the monocular two-dimensional human body pose estimation network in sequence to obtain the multi-view two-dimensional human body poses under multi-view images;

[0078] Step 1.3: Generate multiple left and right image pairs. Denote the images collected by the multi-view camera as C1, C2, C3... Cn, where n is the number of multi-view cameras, and take every two adjacent multi-view cameras as a group to obtain multiple groups of binocular image pairs {C1, C2}, {C2, C3}... {Cn - 1, Cn}. Denote the image with the smaller serial number in each group as the left image, and the other image as the right image, to form n - 1 left and right image pairs. Use the two-dimensional human body pose estimation results of each group of left and right image pairs as the input for the subsequent steps;

[0079] Step 1.4: Generate multiple candidate 3D human pose estimation results. For the input of the 2D human pose estimation results of each pair of left and right images, denoted as (2, 2, N), which represents the (x, y) coordinates of N human skeleton key points in the 2D image coordinate system of 2 left and right images. The first 2 represents 2 images, the second 2 represents the two dimensions of length and width in the image pixel coordinate system, and N represents the number of key points of the human skeleton. This number is set by the ground truth of the human pose dataset used in training. Since in the triangulation method, a small error in the 2D human pose estimation results will cause the 3D coordinate points calculated by the triangulation method to have coordinate offsets in the direction perpendicular to the binocular camera baseline, significantly affecting the accuracy of 3D human pose estimation. Therefore, in this algorithm, a coordinate offset along the binocular epipolar line direction is given to the 2D human pose estimation coordinate points of the right image. The offset amount is generally 1 pixel unit length, and when the baseline is short, it is generally 0.5 pixel unit length. The offset amount is a hyperparameter and is manually set by humans. Based on the experience summarized from comparing the effects of different set offset amounts in multiple repeated experiments, multiple possible 2D human pose estimation coordinate results of the right image are obtained. For example, the coordinate point offset length is d pixels, and the number of offset times is m - 1 times. Finally, multiple candidate results of the 2D human pose estimation coordinates of the right image in the (1, 2, N, m) matrix data format are obtained. Among them, (1, 2, N) is a 2D human pose estimation result, and m - 1 groups are generated through the above process. Adding the initial group, there are a total of m groups of 2D human pose estimation candidate results. Through experiments, it is proved that setting offsets for both the left and right image 2D human pose key points simultaneously has an approximate effect to setting offsets only for the right image key points, and will significantly increase the calculation amount. Therefore, only the 2D human key point coordinates of the right image are set with offsets.

[0080] Given the camera parameters, fix the 2D human pose estimation results of the left image, and calculate the 3D human pose estimation coordinate matrix data format of (3, N, m) through the triangulation method with m possible 2D human key point coordinates of the right image respectively, representing the 3D coordinate values of N skeleton key points of the measured human in m candidate world coordinate systems; among them, 3 represents the three dimensions of x, y, and z of the 3D coordinate values in the world coordinate system.

[0081] Specifically, the said Step 2 includes:

[0082] Step 2.1: Calculate the difference in the lengths of symmetric limbs. According to the prior knowledge of human skeleton lengths, the estimated results of symmetric human limbs should have the same length, generally including limbs such as the left and right forearms, upper arms, left and right calves, and left and right thighs. Therefore, calculate the symmetric limb length error for the 3D key point coordinates of each group of 3D human pose estimation results as follows:

[0083]

[0084] Among them, R is the set of limbs with symmetry in the three-dimensional human skeleton, x1 and x2 are any pair of limbs with symmetry in this set, L is the limb length calculated from the coordinates of the key points of the two ends of the limb in the three-dimensional Euclidean coordinate system, and abs represents the absolute value.

[0085] Step 2.2: Calculate the difference between the limb length ratio and the prior. According to the prior of the human limb length ratio, although the limb length ratios of different people are somewhat different, they are all within a certain range. The variation range of the ratio of each human limb to the length of the human torso, which is used as the reference limb, is denoted as the prior range of the human limb length ratio. This prior is obtained by statistically analyzing the true three-dimensional human postures of multiple collectors of different genders, ages, and body shapes in multiple public datasets and actual scene data collection. The initial value includes the mean and variance of the ratio of each human limb length to the torso length, and a reasonable range of the prior of the human limb length ratio is constructed based on the Gaussian distribution. For multiple candidate three-dimensional human pose estimation results, calculate their limb length ratios respectively, and further calculate the absolute difference between the limb length ratio and the reasonable value of the prior of the human limb length ratio. The sum of the differences obtained for all limbs is denoted as the prior error of the limb length ratio as follows:

[0086]

[0087] Among them, G is the set of limbs of the three-dimensional human skeleton except for the upper body torso of the human body used as the reference, x is any limb in this set, y is the torso limb of the human body with L as the reference, L is the limb length calculated from the coordinates of the key points of the two ends of each limb in the three-dimensional Euclidean coordinate system, P represents the mean of the prior of the x limb in the human limb length ratio, abs represents the absolute value, k is a weight hyperparameter, which is fixed in the experiment. If the value within abs exceeds twice the variance of the prior of the x limb in the human limb length ratio, then k is 1, otherwise it is 0.5.

[0088] Step 2.3: Calculate the radian value of the included angle between connected limbs exceeding the prior range. According to the prior of the human skeleton, each human limb has biological movement angle limitations on the human skeleton, which is denoted as the prior of the human movement angle. This prior is obtained by statistically analyzing the true three-dimensional human postures of multiple collectors of different genders, ages, and body shapes in multiple public datasets and actual scene data collection, including the movement boundaries and reachable areas of human joint angles. For multiple candidate three-dimensional human pose estimation results, calculate the radian values of their joint movement angles respectively, and take the absolute difference with the radian values of the prior range of human kinematics. The sum of the differences obtained for all limbs is denoted as the kinematic prior error as follows:

[0089]

[0090] Where r is the radian value of the angle calculated in the three-dimensional Euclidean space coordinate system for two limbs connected by any three-dimensional human bone key point, V is the range of the human motion angle corresponding to this human bone key point in the prior of human motion angles, V min and V max represent two radian edge values of V, one large and one small, and min() is the operation of taking the smaller value.

[0091] Step 2.4: Calculate the rationality error score of the three-dimensional pose.

[0092] The final rationality error score of the three-dimensional pose is the weighted sum of the scores of the three parts in Steps 2.1, 2.2, and 2.3, denoted as:

[0093] Loss = α × Loss1 + b × Loss2 + c × Loss3

[0094] where a, b, and c are the weighted hyperparameters of the scores of the three parts.

[0095] Step 2.5: Calculate the rationality error scores of the three-dimensional poses for multiple candidate three-dimensional human poses respectively, and finally select the three-dimensional human pose corresponding to the smallest rationality error score of the three-dimensional pose as the optimized reconstruction result of the three-dimensional human pose for this pair of left and right images.

[0096] Specifically, the said Step 3 includes:

[0097] For a multi-view system, Step 1 refines the multi-view system into a combination of multiple binocular systems, and performs the operations of Step 1 and Step 2 under each group of binocular systems. Finally, the coordinates of the three-dimensional human pose key points under each group of binocular systems are obtained. And these multiple groups of binocular systems essentially correspond to the same real three-dimensional human pose. Therefore, it is necessary to fuse the optimized reconstruction results of the three-dimensional human poses under multiple binocular systems to obtain the final unique optimized reconstruction result of the three-dimensional human pose. The fusion criterion is the reprojection error of the optimized reconstruction result of the three-dimensional human pose of each group of binocular systems in other groups of binocular systems. The three-dimensional human pose points of the group of binocular systems with smaller errors have greater weights in the weighted average, that is, the softmax weighted average method. The final unique optimized reconstruction result of the three-dimensional human pose is calculated through the weighted average formula.

[0098] More specifically, for the (3, N, n - 1) groups of three-dimensional human pose coordinate points obtained in step 2, they represent the optimized reconstruction results of the three-dimensional human poses of n - 1 pairs of left and right images. For each three-dimensional human pose key point, calculate the reprojection error of its image coordinates corresponding to the key points of the multi-view two-dimensional human pose in n views respectively, obtaining a data matrix format of (3, N, (n - 1)×n), and perform inverse normalization and weighted average on the multi-view camera reprojection error dimension, that is, the smaller the reprojection error of the three-dimensional human pose key point coordinates, the greater the weight in the weighted average. Finally, obtain the optimized reconstruction result of the multi-view three-dimensional human pose of (3, N).

[0099] The present invention also provides a multi-view three-dimensional human pose optimization and reconstruction system based on human skeleton prior. The multi-view three-dimensional human pose optimization and reconstruction system based on human skeleton prior can be implemented by executing the process steps of the multi-view three-dimensional human pose optimization and reconstruction method based on human skeleton prior. That is, those skilled in the art can understand the multi-view three-dimensional human pose optimization and reconstruction method based on human skeleton prior as a preferred implementation manner of the multi-view three-dimensional human pose optimization and reconstruction system based on human skeleton prior.

[0100] Embodiment 2

[0101] Embodiment 2 is a preferred example of Embodiment 1

[0102] According to a multi-view three-dimensional human pose optimization and reconstruction method provided by the present invention, it includes:

[0103] Collect images based on multi-view synchronous cameras, and use an open-source algorithm with Resnet152 as the backbone network to give the preliminary result of two-dimensional human pose estimation in multi-view images as the input.

[0104] Based on the publicly available Human3.6m dataset and MPII dataset, conduct statistics on the prior of human skeleton length ratio and motion angle. When conducting actual data tests, set the number of multi-view cameras to 4, and set the weight hyperparameters a, b, and c to 0.5, 0.3, and 0.2 respectively when constructing the Loss, meeting the practical requirement that the average error of the full-body degrees of freedom is less than 5 cm in the actual indoor data test.

[0105] Furthermore, in this embodiment, the multi-view three-dimensional human pose reconstruction network based on deep learning can, according to actual application needs, use other network models that can output the initial three-dimensional human pose reconstruction results to achieve the effects of reducing the resources required for calculation and accelerating the calculation speed;

[0106] In this embodiment, the human skeleton prior optimization method can be adjusted according to the number and positions of key points of a specific pose model, and based on the actual physical meanings of the key points of the pose model, the effects of meeting the prior knowledge of human bone length, symmetry, and kinematics can be achieved;

[0107] In this embodiment, the human skeleton prior optimization method can be adjusted according to the representation form of the three-dimensional space pose during the actual use process to meet the requirements of general post-processing methods such as temporal smoothing in the actual scenario.

[0108] Those skilled in the art know that in addition to implementing the systems, devices, and their respective modules provided by the present invention in the form of pure computer-readable program codes, the method steps can be logically programmed to enable the systems, devices, and their respective modules provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers, etc., to achieve the same program. Therefore, the systems, devices, and their respective modules provided by the present invention can be regarded as a kind of hardware component, and the modules included therein for implementing various programs can also be regarded as the structures within the hardware component; the modules for implementing various functions can also be regarded as either software programs for implementing the methods or the structures within the hardware component.

[0109] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined arbitrarily with each other.

Claims

1. A multi-view three-dimensional human pose optimization and reconstruction method based on human skeleton prior, characterized in that Including: Step S1: Obtain a multi-view human body image, and obtain a multi-view two-dimensional human body pose estimation result based on the obtained multi-view human body image; Group the obtained multi-view two-dimensional human body pose estimation results in pairs, and generate multiple candidate three-dimensional human body pose reconstruction results under binocular vision through triangulation calculation and introducing coordinate point offset; Step S2: Based on multiple candidate three-dimensional human body pose reconstruction results under binocular vision, calculate the three-dimensional pose rationality error score through the human bone prior optimization rule, and select the three-dimensional human body pose with the smallest error score as the optimized result of the three-dimensional human body pose reconstruction under multiple binocular visions; Step S3: Based on the optimized results of the three-dimensional human body pose estimation under multiple binocular visions, optimize to obtain the final three-dimensional human body pose reconstruction result through the multi-view reprojection error minimization method.

2. The multi-view three-dimensional human pose optimization and reconstruction method based on human skeleton prior according to claim 1, wherein The said Step S1 includes: Step S1.1: Obtain a multi-view human body image based on the acquisition method of synchronization and calibration by any multi-view camera; Step S1.2: Input the multi-view human body image into the monocular two-dimensional human body pose estimation network in sequence to obtain the multi-view two-dimensional human body pose under the multi-view human body image; Record the multi-view human body images as C1, C2, C3... Cn, where n is the number of multi-view cameras, and take every two adjacent multi-view cameras as a group to obtain multiple pairs of binocular images {C1, C2}, {C2, C3}... {Cn-1, Cn}, record the image with the smaller serial number in each group as the left image, and the other image as the right image to form n-1 left-right image pairs; Step S1.4: For the two-dimensional human body pose estimation result of the right image in the two-dimensional human body pose of each group of left-right image pairs, perform coordinate point offset along the binocular epipolar line direction that meets the preset requirements to obtain multiple two-dimensional human body pose estimation coordinate results of the right image; Step S1.5: Based on the fixed two-dimensional human body pose estimation result of the left image, calculate multiple three-dimensional human body pose estimation results through triangulation method respectively with multiple two-dimensional human body pose estimation coordinate results of the right image.

3. The multi-view three-dimensional human pose optimization and reconstruction method based on human skeleton prior according to claim 1, characterized in that The said Step S2 includes: Step S2.1: According to the prior of human bone length, the estimated results of symmetric human limbs should have the same length, and calculate the symmetric limb length error for the three-dimensional key point coordinates of each three-dimensional human body pose estimation result: where R is the set of limbs with symmetry in the three-dimensional human skeleton, x1, x2 are any pair of symmetric limbs in this set, L is the limb length calculated from the coordinates of the two end bone key points of this limb in the three-dimensional Euclidean coordinate system, and abs represents the absolute value; Step S2.2: Based on multiple three-dimensional human body pose estimation results, calculate their limb length ratios respectively, and further calculate the absolute difference between the limb length ratio and the reasonable range of the prior of each human limb length ratio, and sum the differences obtained for all limbs, which is recorded as the prior error of the limb length ratio; Among them, G is the set of limbs of the three-dimensional human skeleton except for the upper body trunk of the human body used as a reference. x is any limb in this set, y is the limb of the human body trunk used as a reference for L, L is the limb length calculated from the coordinates of the two end bone key points of each limb in the three-dimensional Euclidean coordinate system, P represents the mean of the prior of the human limb length ratio of the x limb, abs represents the absolute value, and k is the weight hyperparameter; Step S2.3: For multiple three-dimensional human pose estimation results, calculate the radian values of their joint movement angles respectively, and take the absolute difference from the radian values of the prior range of human kinematics. Sum the differences obtained for all limbs, and denote it as the kinematic prior error; Wherein, r is the radian value of the angle calculated in the three-dimensional Euclidean space coordinate system for two limbs connected by any three-dimensional human bone key point, V is the range of human motion angles corresponding to the human bone key point in the prior of human motion angles, V min and V max represent two radian edge values of V, one large and one small, and min() is the operation of taking the smaller value; Step S2.4: Calculate the three-dimensional pose rationality error score: Loss = a×Loss1 + b×Loss2 + c×Loss3 Among them, a, b, and c are the weighted hyperparameters of the scores of the three parts; Step S2.5: Calculate the three-dimensional pose rationality error scores for multiple candidate three-dimensional human poses respectively, and finally select the three-dimensional human pose corresponding to the smallest three-dimensional pose rationality error score as the optimized reconstruction result of the three-dimensional human pose of the left and right image pairs.

4. The multi-view three-dimensional human pose optimization and reconstruction method based on human skeleton prior according to claim 3, wherein The reasonable range of the prior of the length ratio of each human limb includes: Denote the change range of the ratio of each human limb to the length of the human body trunk used as the reference limb as the prior range of the human limb length ratio; Based on the three-dimensional human pose ground truth of multiple acquisition personnel of different genders, ages, and body shapes in multiple public data sets and actual scene acquisition data, the initial values are statistically obtained. The initial values include the mean and variance of the ratio of the length of each human limb to the trunk length, and a reasonable range of the prior of the length ratio of each human limb is constructed based on the Gaussian distribution.

5. The multi-view three-dimensional human pose optimization and reconstruction method based on the human skeleton prior according to claim 3, characterized in that The radian values of the prior range of human kinematics include: According to the prior of the human skeleton, each human limb has biological movement angle limitations on the human skeleton, denoted as the prior of the human movement angle; The prior range of the human movement angle is obtained based on the statistical results of the three-dimensional human pose ground truth of multiple acquisition personnel of different genders, ages, and body shapes in multiple public data sets and actual scene acquisition data, and includes the movement boundaries and reachable regions of human joint angles.

6. The multi-view three-dimensional human pose optimization and reconstruction method based on the prior of human skeleton according to claim 1, wherein, The step S3 includes: Based on the optimized results of three-dimensional human pose estimation under multiple groups of binocular vision, calculate the reprojection error of each optimized result of three-dimensional human pose estimation under binocular vision in other groups of binocular vision respectively, and calculate the final optimized reconstruction result of multi-view three-dimensional human pose through the softmax weighted average method.

7. A multi-view three-dimensional human pose optimization and reconstruction method based on human skeleton prior, characterized in that, Including: Module M1: Obtain multi-view human body images, and obtain multi-view two-dimensional human pose estimation results based on the obtained multi-view human body images; Group the obtained multi-view two-dimensional human pose estimation results in pairs, and generate candidate three-dimensional human pose reconstruction results under multiple groups of binocular vision by calculating with the triangulation method and introducing coordinate point offsets; Module M2: Based on the candidate three-dimensional human pose reconstruction results under multiple groups of binocular vision, calculate the three-dimensional pose rationality error score through the human bone prior optimization rule, and select the three-dimensional human pose with the smallest error score as the optimized reconstruction result of the three-dimensional human pose under multiple groups of binocular vision; Module M3: Based on the optimization results of three-dimensional human pose estimation under multiple binocular visions, the final three-dimensional human pose reconstruction result is optimized by the method of minimizing the multi-view reprojection error.

8. The multi-view three-dimensional human pose optimization and reconstruction system based on the prior of the human skeleton according to claim 7, wherein, The said module M1 includes: Module M1.1: Obtain multi-view human images through the acquisition method of synchronization and calibration of any multi-view cameras; Module M1.2: Input the multi-view human images into the monocular two-dimensional human pose estimation network in sequence to obtain the multi-view two-dimensional human poses under the multi-view human images; Denote the multi-view human images as C1, C2, C3... Cn, where n is the number of multi-view cameras, and take every two adjacent multi-view cameras as a group to obtain multiple pairs of binocular images {C1, C2}, {C2, C3}... {Cn-1, Cn}. Denote the image with a smaller serial number in each group as the left image, and the other image as the right image to form n - 1 pairs of left and right images; For the two-dimensional human pose estimation result of the right image in the two-dimensional human pose of each pair of left and right images, perform coordinate point offset along the binocular epipolar line direction that meets the preset requirements to obtain multiple two-dimensional human pose estimation coordinate results of the right image; Based on the fixed two-dimensional human pose estimation result of the left image, calculate multiple three-dimensional human pose estimation results respectively through triangulation with multiple two-dimensional human pose estimation coordinate results of the right image.

9. The multi-view three-dimensional human pose optimization and reconstruction system based on prior knowledge of human skeletons according to claim 7, wherein, The said module M2 includes: Module M2.1: According to the prior of human bone length, the estimated results of symmetric human limbs should have the same length, and calculate the symmetric limb length error for the three-dimensional key point coordinates of each three-dimensional human pose estimation result: where R is the set of limbs with symmetry in the three-dimensional human skeleton, x1, x2 are any pair of symmetric limbs in this set, L is the limb length calculated from the bone key point coordinates at both ends of the limb in the three-dimensional Euclidean coordinate system, and abs represents the absolute value; Module M2.2: Based on multiple three-dimensional human pose estimation results, calculate their limb length ratios respectively, and further calculate the absolute difference between the limb length ratio and the reasonable range of the prior of human limb length ratios. Sum the differences obtained for all limbs and denote it as the prior error of limb length ratio; where G is the set of limbs of the three-dimensional human skeleton except for the human upper body trunk used as the reference, x is any limb in this set, y is the human trunk limb with L as the reference, L is the limb length calculated from the bone key point coordinates at both ends of each limb in the three-dimensional Euclidean coordinate system, P represents the mean of the prior of the human limb length ratio of limb x, abs represents the absolute value, and k is the weight hyperparameter; Module M2.3: For multiple three-dimensional human pose estimation results, calculate the radian values of the joint movement angles respectively, and take the absolute difference with the radian values in the prior range of human kinematics. Sum the differences obtained for all limbs and denote it as the prior error of kinematics; Wherein, r is the radian value of the angle calculated in the three-dimensional Euclidean space coordinate system for two limbs connected by any three-dimensional human body bone key point, V is the range of the human body movement angle corresponding to the human body bone key point in the prior of the human body movement angle, V min and V max represent two radian edge values of V, one large and one small, and min() is the operation of taking the smaller value; Module M2.4: Calculate the rationality error score of the three-dimensional pose: Loss = α × Loss1 + b × Loss2 + c × Loss3 where a, b, c are the weighted hyperparameters of the scores of the three parts; Module M2.5: Calculate the 3D pose rationality error scores for multiple candidate 3D human poses respectively, and finally select the 3D human pose corresponding to the smallest 3D pose rationality error score as the optimized reconstruction result of the 3D human pose for the left and right image pairs.

10. The multi-view three-dimensional human pose optimization and reconstruction system based on human skeleton prior according to claim 7, characterized in that, The module M3 includes: Based on the optimized results of 3D human pose estimation under multiple groups of binocular vision, calculate the reprojection error of each group of optimized results of 3D human pose estimation under binocular vision in other groups of binocular vision respectively, and calculate the final optimized reconstruction result of multi-view 3D human pose through the softmax weighted average method.

Citation Information

Patent Citations

  • Human body three-dimensional posture reconstruction method and system based on multi-view human body image

    CN116152439A