Three-dimensional human pose reconstruction method based on multi-camera self-calibration technology
By using multi-camera self-calibration technology, camera pose and parameters are optimized using camera arrays and SMPL models, solving the problems of occlusion and calibration complexity in traditional human pose estimation, and achieving high-precision and robust human pose reconstruction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SUN YAT SEN UNIV
- Filing Date
- 2023-12-06
- Publication Date
- 2026-05-19
AI Technical Summary
Traditional single-camera-based human pose estimation is susceptible to occlusion and tracking failures, while multi-camera system calibration is cumbersome and its accuracy is easily affected. Existing non-contact measurement devices lack accuracy and robustness in human motion capture.
A multi-camera self-calibration technique is adopted, which acquires images in real time through a camera array, uses the OpenPose detector and MAGSAC++ estimator to obtain 2D human joints, and combines triangulation and the Perspective-n-Point method to iteratively optimize the camera pose and 3D joints, constructs an SMPL model, uses the length of the limb bones as a scale factor to optimize the camera calibration relationship, and iteratively adjusts the morphology and pose parameters.
It eliminates the need for cumbersome external calibration, improves the accuracy and robustness of human posture measurement, solves the problems of calibration complexity and motion capture error accumulation in multi-camera systems, and provides more accurate human posture reconstruction.
Smart Images

Figure CN117635843B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of three-dimensional human body reconstruction technology, and in particular to a three-dimensional human body pose reconstruction method based on multi-camera self-calibration technology. Background Technology
[0002] Based on the measurement method, 3D human pose reconstruction can be divided into two types: contact measurement and non-contact measurement. Contact measurement requires attaching special markers to the human body for motion capture. Its single-point tracking measurement accuracy is high, reaching the millimeter level; however, contact measurement equipment is complex and bulky, making it inconvenient to use. Furthermore, attaching markers to the human body can interfere with normal human movement. Non-contact measurement uses a single camera or a measurement network composed of multiple cameras. It does not require the person being measured to wear any equipment and has advantages such as small size, light weight, and low power consumption. However, traditional single-camera-based human pose estimation methods are prone to occlusion and tracking failures, and the measurement accuracy and algorithm robustness are often unsatisfactory. Multi-camera system-based measurement methods can better solve problems such as occlusion and depth blur, but before deploying a multi-camera system, tedious manual initialization work such as camera calibration is usually required, and the final measurement accuracy is easily affected by the camera calibration results and the range of human pose movement. Summary of the Invention
[0003] Therefore, it is necessary to provide a more effective and efficient motion capture method that does not require cumbersome external calibration to address the above-mentioned technical problems, thereby improving the accuracy and robustness of human posture measurement and the 3D human posture reconstruction method based on multi-camera self-calibration technology.
[0004] This invention provides a method for three-dimensional human pose reconstruction based on multi-camera self-calibration technology, the method comprising:
[0005] Construct a camera array for real-time acquisition of human motion sequences;
[0006] Based on the images acquired by the camera array, the initial pose transformation between the two cameras and the coordinate information of the 3D human joint points are obtained.
[0007] An SMPL model is constructed based on 3D human body scan data. The lengths of the human limb bones are obtained based on the SMPL model. The SMPL model describes the human posture through morphological parameters and posture parameters.
[0008] A new camera is added to the camera array. Based on the length of the limb bones and the 3D human joint points obtained from the intersection of the initial two cameras and the 2D human joint points under the new camera, the calibration relationship between the new camera and the camera array is continuously optimized iteratively.
[0009] Based on the camera array after the addition of the camera, the initial 3D human body joints are optimized to obtain optimized 3D human body joints. The distance between human body joints in the optimized 3D human body joints is calculated to calculate the morphological parameters and obtain the optimal morphological parameters.
[0010] Based on the changes in the relative rotation angles of the 3D human joints, the posture parameters in the SMPL model are recovered to obtain the current posture parameters.
[0011] The three-dimensional human posture is reconstructed based on the morphological parameters and the current posture parameters.
[0012] Furthermore, the step of obtaining the initial pose transformation between the two cameras and the 3D human joint coordinate information based on the images acquired by the camera array includes:
[0013] Based on the images acquired by the camera array, the OpenPose detector is used to estimate the pose of the human body in the images, and outputs the two-dimensional coordinates and confidence scores of a set number of 2D human body joints.
[0014] Based on the 2D human joints, the fundamental matrix is estimated using the MAGSAC++ estimator, and the initial camera pose is obtained by decomposition.
[0015] Based on the 2D human joints and the initial camera pose, depth information is recovered using triangulation to recover the corresponding spatial 3D human joint nodes.
[0016] Based on the 3D human joint nodes and the initial camera pose, the remaining cameras are added sequentially. Based on the Perspective-n-Point method corresponding to 2D-3D points, the camera pose and the extraction results of 3D human joint nodes are iteratively optimized. Combined with the loss function, the iterative optimization process of solving the camera pose and the extraction results of 3D joint nodes are further described in detail.
[0017] Furthermore, the construction of the SMPL model based on the 3D human body scan data includes:
[0018] The human body is scanned in three dimensions using a scanner, resulting in an unstructured disordered point cloud. The FARM algorithm is then used to construct a personalized SMPL model from the scanned unstructured disordered point cloud.
[0019] The SMPL model includes morphological parameters and posture parameters. The morphological parameters describe the human body's physical characteristics by setting values for a number of dimensions, while the posture parameters describe the human body's motion posture at a certain moment by setting the relative rotation angles of a number of joints.
[0020] Furthermore, based on the SMPL model, the lengths of the human limb bones are obtained; wherein, the SMPL model describes the human posture through morphological parameters and pose parameters, and the SMPL model is as follows:
[0021]
[0022]
[0023] in, For morphological parameters, These are attitude parameters. A posture indicating a still body. This represents the default template for the SMPL model. and These are vertex vectors representing template offsets generated by morphological and pose parameters, respectively. Represents the coordinates of the three-dimensional joints in the SMPL model; For linear hybrid skinning functions, it means that by adjusting the attitude parameters, one can achieve the desired effect. Different actions are presented on it.
[0024] Furthermore, a new camera is added to the camera array. Based on the length of the limb bones and the 3D human joint points obtained from the initial dual-camera intersection with the 2D human joint points under the new camera, the calibration relationship between the new camera and the camera array is iteratively optimized. Based on the camera array after the addition of the new camera, the initial 3D human joint points are optimized to obtain optimized 3D human joint points. The distance between human joint points in the optimized 3D human joint points is calculated to optimize the morphological parameters. The calculation of the morphological parameters includes:
[0025] Add other cameras, iteratively optimize the 3D human joints obtained by triangulation, and calculate the distance between human joints to obtain the optimal solution of morphological parameters. The loss function is:
[0026]
[0027] in, This indicates the error caused by the joints of the 3D human body. λ represents the discrete point cloud error. bone and λ prior They represent and The weights;
[0028] Errors caused by 3D human joints are calculated by taking the length of the limb bones obtained from the joint vertices in the SMPL model. The length B of the limb skeleton obtained by triangulating the 3D human joints using a multi-camera system iThe optimization of minimizing the Euclidean distance is as follows:
[0029]
[0030] Furthermore, it also includes denoising the morphological parameters: denoising is performed using Gaussian filtering, with the loss function being;
[0031]
[0032] Furthermore, obtaining the current pose parameters in the SMPL model based on the changes in the relative rotation angles of the 3D human joints includes:
[0033] Based on the changes in the relative rotation angles of the 3D human joints, pose parameters can be recovered using the SMPL model:
[0034] Its attitude energy function can be expressed as:
[0035]
[0036] in, This represents the error of two-dimensional joints in the human body. This represents the error in predicting the 3D joint points from the intersection of 2D human joint points under multiple views. λ represents the discrete point cloud error. j2d , λ j3d and λ prior Corresponding to and The weights;
[0037] The two-dimensional coordinates of the human body joints are obtained by minimizing the projection of the three-dimensional joint coordinates of the adjusted SMPL model. With the established personalized two-dimensional joint coordinates of the human body The Euclidean distance between them is optimized as follows:
[0038]
[0039] The 3D joint coordinates of the human body can be minimized by adjusting the 3D joint coordinates of the SMPL model. 3D joint coordinates predicted by multi-view intersection The Euclidean distance between them is optimized as follows:
[0040]
[0041] Furthermore, it also includes denoising the attitude parameters: denoising is performed using Gaussian filtering, with the loss function being;
[0042]
[0043] The aforementioned 3D human pose reconstruction method based on a multi-camera self-calibration system, by adding cameras to the camera array, iteratively optimizes the calibration relationship between the new cameras and the camera array based on the limb bone length and the 3D human joints obtained from the initial intersection of the two cameras and the 2D human joints under the new cameras, without the need for cumbersome external calibration of the multi-camera system. The limb bone length is used as a scale factor to recover the true distance between each camera, and the morphological and pose parameters in the SMPL model are adjusted based on the optimized 3D human joints obtained after the addition of the cameras for 3D human pose reconstruction. This solves the problems of complex multi-camera system calibration and error accumulation in human motion capture under multiple views, thereby estimating a more accurate human pose. Attached Figure Description
[0044] Figure 1 This is a flowchart illustrating the method of the present invention in one embodiment;
[0045] Figure 2 This is a schematic diagram of the modules of the method of the present invention in one embodiment;
[0046] Figure 3 This is a schematic diagram illustrating the recovery of depth information of point P based on a triangulation method in one embodiment. Detailed Implementation
[0047] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0048] The 3D human pose reconstruction method based on a multi-camera self-calibration system provided in this application mainly consists of two modules: a multi-camera self-calibration module and a 3D human pose reconstruction module. The specific process is as follows: Figure 1 and Figure 2 As shown.
[0049] The multi-camera self-calibration module performs the following steps:
[0050] S1. Construct a camera array for real-time acquisition of human motion sequences to obtain color images of the human body in real time.
[0051] In one embodiment, in order to obtain human posture information from multiple perspectives, a camera array consisting of four color cameras is constructed. The construction site can be inside a space station or an experimental site simulating a human working scenario in orbit, so as to collect the motion sequence of the human body in real time during the performance of the task.
[0052] S2. Based on the color image, obtain 2D human joint coordinate information and skeleton using the OpenPose detector;
[0053] In one embodiment, based on the color image, the OpenPose detector is used to estimate the pose of the human body in the image, and outputs the two-dimensional coordinates and confidence scores of 25 human body joints.
[0054] The 25 human joint points include: head, neck, left shoulder, left elbow, left wrist, left palm, left thumb, left middle fingertip, right shoulder, right elbow, right wrist, right palm, right thumb, right middle fingertip, upper spine, middle spine, lower spine, left hip, left knee, left ankle, left foot, right hip, right knee, right ankle, and right foot;
[0055] The confidence level refers to a 2D representation of the belief that the 25 human joints appear at each pixel location. The confidence level of a pixel at a joint is 1, and the surrounding pixels are distributed in a Gaussian pattern according to distance; the closer the distance, the higher the confidence level, and vice versa.
[0056] S3. Based on the 2D human joints, estimate the fundamental matrix using the MAGSAC++ estimator and decompose it to obtain the initial values of the camera extrinsic parameters.
[0057] The image points of the same 3D point under different viewpoints exhibit epipolar constraints, and the fundamental matrix is the algebraic representation of these constraints. This epipolar constraint relationship is independent of the scene structure and depends only on the camera's intrinsic and extrinsic parameters (relative pose). Therefore, by matching 2D human joints, the fundamental and essential matrices of the two images are estimated. Then, singular value decomposition is performed on the essential matrix to obtain the initial relative pose of the camera, i.e., the rotation matrix R and translation vector t of camera two relative to camera one, as described in detail below:
[0058] The fundamental matrix refers to the mapping from image point p1 in one image to the epipolar line l2 in another image, which can be represented as l2 = Fp1. The other image point p2 that matches image point p1 must lie on the epipolar line l2, hence... The fundamental matrix F is a 3×3 matrix with 9 unknowns. However, the homogeneous coordinates used in the formula above are equal when they differ by a constant factor, meaning that only n sets (n≥8) of matching points need to be given. A fundamental matrix F that satisfies the epipolar constraint can then be calculated.
[0059] In one embodiment, since the human body in the target space can provide 25 sets of matching point pairs coordinates, and each set of matching points has different magnitudes of error, the MAGSAC++ estimator is used to select 8 sets of matching point pairs with the smallest error under epipolar constraints to improve the accuracy of estimating the fundamental matrix F, making the final decomposed camera relative pose more reliable.
[0060] MAGSAC++ is a fast, reliable, accurate, and robust estimator. It uses edge sample consensus, estimates the model threshold using the noise parameter σ, replaces the original least squares with weighted least squares, and calculates weights through a marginalization process. Therefore, its sensitivity to setting thresholds for inlier-outlier pairs is significantly lower than other robust estimators. The MAGSAC++ estimator can iteratively estimate parameters that conform to the mathematical model of a set of outlier datasets (here referring to human joint point pairs that cause significant matching errors).
[0061] To improve the stability and accuracy of the solution, the coordinates of the 25 sets of human joint points input were first normalized. Isotropic normalization was then applied based on... Transform the image coordinates. Here, T and T′ are normalized transformations consisting of translation and scaling.
[0062] For each possible value of σ, the probability that point p belongs to the interior set P is:
[0063]
[0064] Where σ is the noise parameter, C(n) is the contrast ratio, and D(θ) is the noise parameter. σ ,p) represents the error from the point to the model.
[0065]
[0066] The weight of point p is:
[0067] ω(D(θ i ,p))=∫P(p|θ i ,σ)f(σ)dσ
[0068] The weighting function is a marginal function of the inlier error:
[0069] ω(r)=∫g(r|σ)f(σ)dσ
[0070] Based on the eight selected pairs of matching points with the smallest epipolar constraint errors, the estimated fundamental matrix F can be obtained by solving the linear equation, and then the essential matrix E can be solved, as shown in the following equation:
[0071] E=C T FC
[0072] Where C is the known intrinsic parameter matrix of the camera.
[0073] Singular value decomposition (SVD) is performed on the obtained essential matrix. Its singular values are in the form of [σ,σ,0]. This yields the rotation matrix R and translation vector t of camera 2 relative to camera 1, thus obtaining the initial pose transformation between the two cameras.
[0074] S4. Based on the 2D human joints and the initial camera pose, recover the depth information using the triangulation method, and recover the corresponding spatial 3D human joint nodes, thus obtaining the initial 3D human joint coordinate information.
[0075] The triangulation method involves observing the same three-dimensional point P from two different locations, and then using triangulation relationships to reconstruct the depth information of the three-dimensional point P in space, given the coordinates of the two-dimensional projection points of the three-dimensional point observed from these two locations.
[0076] Let the optical centers of camera 1 and camera 2 be O1, O1, and the corresponding point in 3D space for image point P1 and its matching image point P2 be P. From the mapping relationship between the pixel coordinate system and the camera coordinate system, we can obtain:
[0077]
[0078] Where s1 and s2 are the depths of the human joints in the two images, respectively, let x1 = K -1 p1,x2=K -1 Substituting p2 into the above equation, we get:
[0079] s2x2=s1Px1+t
[0080] Multiply both sides of the equation by an antisymmetric matrix on the left. have to:
[0081]
[0082] Finally, by substituting s1 back into the system of equations relating pixel coordinates to camera coordinates, s2 can be obtained, thus acquiring the depth information of point P.
[0083] However, due to problems such as camera calibration errors and inaccurate estimation of human joint positions during the experiment, the two rays will not intersect at a single point in space, such as... Figure 3 As shown.
[0084] In one embodiment, to minimize the object error, the following loss function is set:
[0085]
[0086] Where C represents the number of cameras, ω i This represents the confidence level of the key points output by OpenPose from the image under camera i. po i This represents the distance from point P to the optical center of the camera.
[0087] S5. Based on the 3D human joint nodes and the initial camera pose, add the remaining cameras in sequence, and iteratively optimize the solution of camera pose and 3D human joint point extraction results based on the PnP (Perspective-n-Point) method corresponding to 2D-3D points.
[0088] In one embodiment, 25 3D human body joint points in the world coordinate system (the coordinates of the i-th 3D point are marked) obtained by the intersection of camera one and camera two are used. Matching point pairs of 2D human joints in the coordinate systems of the remaining camera images (the i-th 2D point coordinates are marked). This constitutes a 2D-3D point pair matching, while taking the association weight of each point pair (the association weight of the i-th point pair is denoted as...). The Perspective-n-Point (PnP) problem, i.e., the pose transformation relationship {R|t} between camera coordinate systems, is constructed as a nonlinear least squares problem concerning the confidence of human joints and the reprojection error. Nonlinear optimization is performed using the Levenberg-Marquarelt (LM) algorithm, with the loss function as follows:
[0089]
[0090] Where, ω i,j This represents the confidence level of the human joint points output by OpenPose based on the combined images from camera i and camera j. This represents the reprojection error under camera i, which is the error obtained by comparing the pixel coordinates (the observed projected position) with the position of the 3D point obtained by projecting it according to the currently estimated camera pose.
[0091] The optimal camera pose and 3D human pose are eventually obtained through convergence.
[0092] The 3D human pose reconstruction module performs the following steps:
[0093] S6. Establish a personalized human SMPL model based on the FRAM algorithm;
[0094] The human body of the subject is scanned in 3D using a specific scanner, but the resulting point cloud is unordered and cannot be directly used for 3D human reconstruction. Therefore, by using the FARM algorithm to construct a personalized SMPL model from the scanned unstructured, unordered point cloud, the unordered point cloud can be ordered, and point cloud locations with the same area can be merged, providing crucial data support for subsequent 3D human reconstruction based on the SMPL model.
[0095] Compared to traditional linear blend skinning (LBS) methods, the Skinned Multi-Person Linear Model (SMPL) proposes a method for describing the surface morphology of human pose images, effectively avoiding surface distortion during human model movement. In one embodiment, the SMPL model is as follows:
[0096]
[0097]
[0098] in, Shape parameters can describe human body characteristics through 10 dimensions of values, such as height, weight, and the proportions of different body parts. Pose parameters can be used to describe the human body's motion posture at a certain moment through the relative rotation angles of 24 joints.
[0099] in, The posture indicating a still body, also known as a T-pose. This represents the default template for the SMPL model. and These represent vertex vectors that represent the template offsets generated by the morphological parameters and the pose parameters, respectively. The Linear Hybrid Skin Function (LBS) indicates that different actions can be presented on a T-pose by adjusting the pose parameters.
[0100] S7. Using the length of the limb bones as a scale factor, the true distance between each camera is recovered. Based on the iteratively optimized 3D human joints, the distance between the human joints is calculated to adjust the morphological parameters of the SMPL model. Perform calculations;
[0101] Based on the length of the limb bones and the 3D human joints obtained from the initial intersection of the two cameras and the 2D human joints under the new camera, the calibration relationship between the new camera and the camera array is continuously optimized iteratively.
[0102] In a multi-camera vision system, the ratio of translation vectors between cameras remains consistent with that between camera one and camera two when introducing additional available cameras. After the multi-camera system performs external calibration based on the presented human body, the human pose can be recovered using a pre-scanned mesh model, with almost no change to the bone length of the human limbs. Therefore, the bone length of the limbs of the human body model scanned on the ground can be used as a scale factor to recover the true distance between the cameras.
[0103] Based on the camera array after the addition of the camera, the initial 3D human body joints are optimized to obtain optimized 3D human body joints. The distance between human body joints in the optimized 3D human body joints is calculated to optimize the morphological parameters.
[0104] Morphological parameters By adding additional cameras, the 3D human body joints obtained through triangulation can be continuously optimized iteratively, and the distances between the human body joints can be calculated to obtain the desired results. The optimal solution has the following loss function:
[0105]
[0106] in, This indicates the error caused by the joints of a three-dimensional human body. λ represents the discrete point cloud error. bine and λ prior They represent and The weights;
[0107] Errors caused by 3D human joints can be corrected by calculating the limb bone lengths from the joint vertices in the SMPL model. The length B of the limb skeleton obtained by triangulating the 3D human joints using a multi-camera system i The optimization is performed by minimizing the Euclidean distance, as shown in the following equation:
[0108]
[0109] During the 3D reconstruction of the human body, some discrete point clouds will appear. These point clouds cannot effectively calculate the bone length between the joints of the limbs and will affect the morphological parameters. There is a significant deviation. Considering the distribution characteristics of the point cloud, Gaussian filtering can be used for noise reduction, and the loss function is as follows;
[0110]
[0111] S8. Based on the changes in the relative rotation angles of the 3D human joints after iterative optimization, the current pose parameters are obtained. The human posture is reconstructed based on the optimal morphological parameters and the current posture parameters.
[0112] To accurately recreate the human body's posture, the error accumulation problem in the 3D human body reconstruction process is solved by minimizing the objective function. The posture energy function can be expressed as:
[0113]
[0114] in, This represents the error of two-dimensional joints in the human body. This represents the error in predicting the 3D joint points from the intersection of 2D human joint points under multiple views. λ represents the discrete point cloud error. j2d , λ j3d and λ prior Corresponding to and The weight.
[0115] The two-dimensional coordinates of human body joints can be obtained by minimizing the projection of the three-dimensional joint coordinates of the adjusted SMPL model. With the established personalized two-dimensional joint coordinates of the human body The Euclidean distance between them is optimized as follows.
[0116]
[0117] The 3D joint coordinates of the human body can be minimized by adjusting the 3D joint coordinates of the SMPL model. 3D joint coordinates predicted by multi-view intersection The Euclidean distance between them is optimized as follows.
[0118]
[0119] During the 3D reconstruction of the human body, some discrete point clouds will appear. These point clouds cannot effectively recover the rotation angles of the joints and will lead to unnatural joint poses. Considering the distribution characteristics of the point clouds, Gaussian filtering can be used for noise reduction to effectively reduce error accumulation. The loss function is as follows.
[0120]
[0121] This invention proposes a 3D human pose reconstruction method based on multi-camera self-calibration technology. Based on a multi-camera system, it utilizes 2D human joints detected by OpenPose and their confidence levels to iteratively calibrate the camera pose and the intersecting 3D human joints, achieving self-calibration of the multi-camera system. This eliminates the constraints on the subject imposed by existing contact motion capture systems and eliminates the need for cumbersome external calibration. It also improves the accuracy of the intersecting 3D human joints, better addressing issues such as occlusion and depth blur. The method uses limb bone length as a scale factor to recover the true distance between cameras. Based on the iteratively optimized 3D human joints, it solves for the optimal morphological and pose parameters for 3D human pose reconstruction, resolving the complexity of the multi-camera system calibration process and the error accumulation problem in human motion capture under multiple views, thus estimating a more accurate human pose.
[0122] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0123] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A method for three-dimensional human pose reconstruction based on multi-camera self-calibration technology, characterized in that, The method includes: Construct a camera array for real-time acquisition of human motion sequences; Based on the images acquired by the camera array, the initial pose transformation between the two cameras and the coordinate information of the 3D human joint points are obtained. An SMPL model is constructed based on 3D human body scan data. The lengths of the human limb bones are obtained based on the SMPL model. The SMPL model describes the human posture through morphological parameters and posture parameters. A new camera is added to the camera array. Based on the length of the limb bones and the 3D human joint points obtained from the intersection of the initial two cameras and the 2D human joint points under the new camera, the calibration relationship between the new camera and the camera array is continuously optimized iteratively. Based on the camera array after the addition of the camera, the initial 3D human body joints are optimized to obtain optimized 3D human body joints. The distance between human body joints in the optimized 3D human body joints is calculated to calculate the morphological parameters and obtain the optimal morphological parameters. Based on the changes in the relative rotation angles of the 3D human joints, the current pose parameters in the SMPL model are obtained; The three-dimensional human posture is reconstructed based on the optimal morphological parameters and the current posture parameters; Based on the SMPL model, the lengths of the human limb bones are obtained; wherein, the SMPL model describes human posture through morphological parameters and posture parameters, and the SMPL model is as follows: in, For morphological parameters, These are attitude parameters. A posture indicating a still body. This represents the default template for the SMPL model. and These are vertex vectors representing template offsets generated by morphological and pose parameters, respectively. Represents the coordinates of the three-dimensional joints in the SMPL model; For linear hybrid skinning functions, it represents the effect achieved by adjusting the attitude parameters. Different actions are presented on it; A new camera is added to the camera array. Based on the length of the limb bones and the 3D human joint points obtained from the initial dual-camera intersection, and the 2D human joint points under the new camera, the calibration relationship between the new camera and the camera array is iteratively optimized. Based on the camera array after the addition of the camera, the initial 3D human joint points are optimized to obtain optimized 3D human joint points. The distance between human joint points in the optimized 3D human joint points is calculated to determine the morphological parameters, including: Add other cameras, iteratively optimize the 3D human joints obtained by triangulation, and calculate the distance between human joints to obtain the optimal solution of morphological parameters. The loss function is: in, This indicates the error caused by the joints of the 3D human body. This represents the error in the discrete point cloud. and They represent and The weights; Errors caused by 3D human joints are calculated by taking the length of the limb bones obtained from the joint vertices in the SMPL model. The length of the limb bones obtained by triangulating the 3D human joints using a multi-camera system. The optimization of minimizing the Euclidean distance is as follows: ; The process of obtaining the pose parameters in the SMPL model based on the changes in the relative rotation angles of the 3D human joints includes: Based on the changes in the relative rotation angles of the 3D human joints, the pose parameters are recovered using the SMPL model: Its attitude energy function is expressed as: in, This represents the error of two-dimensional joints in the human body. This represents the error in predicting the 3D joint points from the intersection of 2D human joint points under multiple views. This represents the error in the discrete point cloud. , and Corresponding to , and The weights; The two-dimensional coordinates of the human body joints are obtained by minimizing the projection of the three-dimensional joint coordinates of the adjusted SMPL model. With the established personalized two-dimensional joint coordinates of the human body The Euclidean distance between them is optimized as follows: The 3D joint coordinates of the human body are obtained by minimizing the adjusted SMPL model. 3D joint coordinates predicted by multi-view intersection The Euclidean distance between them is optimized as follows: 。 2. The method according to claim 1, characterized in that, The step of obtaining the initial pose transformation between the two cameras and the 3D human joint coordinate information based on the images acquired by the camera array includes: Based on the images acquired by the camera array, the OpenPose detector is used to estimate the pose of the human body in the images, and outputs the two-dimensional coordinates and confidence scores of a set number of 2D human body joints. Based on the 2D human joints, the fundamental matrix is estimated using the MAGSAC++ estimator, and the initial camera pose is obtained by decomposition. Based on the 2D human joints and the initial camera pose, depth information is recovered using triangulation to recover the corresponding spatial 3D human joint nodes. Based on the 3D human joint nodes and the initial camera pose, the remaining cameras are added sequentially. Based on the Perspective-n-Point method corresponding to 2D-3D points, the camera pose and the extraction results of 3D human joint nodes are iteratively optimized. Combined with the loss function, the iterative optimization process of solving the camera pose and the extraction results of 3D joint nodes are further described in detail.
3. The method according to claim 1 or 2, characterized in that, The construction of the SMPL model based on 3D human body scan data includes: The human body is scanned in three dimensions using a scanner, resulting in an unstructured disordered point cloud. The FARM algorithm is then used to construct a personalized SMPL model from the scanned unstructured disordered point cloud. The SMPL model includes morphological parameters and posture parameters. The morphological parameters describe the human body's physical characteristics by setting values for a number of dimensions, while the posture parameters describe the human body's motion posture at a certain moment by setting the relative rotation angles of a number of joints.