Post-processing action reconstruction method of 4dhuman action capture technology based on multi-view video acquisition
Through multi-view video acquisition and post-processing network optimization methods, the problem of insufficient data noise and motion smoothness in 4DHuman motion capture technology is solved, high-quality character action reconstruction is achieved, and the accuracy and stability of motion capture is improved. It is suitable for virtual reality and animation production.
Patent Information
- Application Number
- CN202510281563.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-29
AI Technical Summary
There are problems in the existing 4DHuman motion capture technology that lacks data noise and motion smoothness, which affects the accuracy and stability of motion capture.
Multi-view video acquisition combined with post-processing network optimization method is adopted, and after the preliminary pkl file is generated through the 4DHuman human motion estimation network, the shape parameter noise reduction processing is performed, joint data clustering is optimized, and the data is supplemented through error analysis and spline interpolation, and finally the SmoothNet smooth network is used for optimization.
It significantly improves the accuracy and stability of motion capture data, ensures the smoothness of the movement, provides high-quality character action reconstruction, and supports applications in areas such as virtual reality and animation production.
Smart Images

Figure CN120388416A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video motion capture in computer vision, and particularly relates to a post-processing motion reconstruction method for 4D human motion capture technology based on multi-view video acquisition. Background Art
[0002] Parametric human models, such as SMPL (Skinned Multi-Person Linear Model), have been widely used in motion capture. The SMPL model can generate realistic 3D human meshes by parametrically representing the shape and pose of the human body.
[0003] 4DHuman is a leading human motion capture and modeling technology based on the SMPL model. It is based on the HMR structure and combines the phalp tracking technology, enabling the reconstruction and tracking of the human body from monocular videos. It adopts the Transformer-based network architecture HMR 2.0, which allows 4DHuman to handle unconventional poses that were previously difficult to reconstruct from a single image. In addition, 4DHuman also performs well in downstream tasks such as action recognition.
[0004] In the process of in-depth analysis of the SMPL data model, an innovative method was adopted to enhance its data interpretation ability. Specifically, the pkl file generated by 4DHuman was converted into a txt format for more intuitive reading and detailed comparison. This conversion not only facilitates data access and processing but also provides convenience for further data optimization. In the converted txt file, the following key data fields were mainly concerned: global orientation (global_orient), body pose (body_pose), shape parameters (beta(s)), and 3D joint positions (3d_joints). These data fields together constitute a complete description of the human model, covering multiple levels from macro to micro.
[0005] Moreover, it is expected to improve the accuracy of video motion capture by adding viewpoints. On the basis of obtaining data fields from the basic viewpoint, other viewpoints are added for correction and optimization to enhance the generalization ability of the model. This method can reduce data bias caused by a single viewpoint, thereby improving the stability and accuracy of the model under different viewpoints. Summary of the Invention
[0006] Objective of the Invention: Aiming at the problems of data noise and insufficient motion smoothness in the existing 4D Human motion capture technology, the present invention proposes a post-processing network optimization method for the 4D Human motion capture technology based on multi-view video acquisition. By shooting the target person from multiple perspectives and obtaining video clips, after generating a preliminary pkl file using the 4D Human body motion estimation network, the shape parameters are further denoised, the joint data is clustered and optimized, and the data is supplemented through error analysis and spline interpolation. Finally, the SmoothNet smoothing network is used to optimize the motion data. This method not only improves the accuracy and stability of the motion capture data, but also realizes high-quality human motion reconstruction and export through the optimized pkl file, providing efficient and accurate technical support for fields such as virtual reality and animation production.
[0007] Technical Solution: To achieve the objective of the present invention, the present invention proposes a post-processing motion reconstruction method for the 4Dhuman motion capture technology based on multi-view video acquisition. The method includes the following steps:
[0008] Step 1: Shoot the target person from n perspectives, where n≥2, to obtain multiple perspective video clips of the target person, and take one of them as the target perspective and the others as auxiliary perspectives;
[0009] Step 2: Input the above n video clips into the 4D Human body motion estimation network to obtain the corresponding n pkl files;
[0010] Step 3: Extract the shape parameters betas of the human body model of a total of m frames in the pkl file of the target perspective to form a sample matrix with m rows, and perform denoising processing to obtain the optimized pkl file 1_best1.pkl;
[0011] Step 4: Extract the 3D coordinates 3d_joints of 45 joints in m frames from the n pkl files, and unify the coordinate systems after being processed by the global rotation matrix globle_orient of this frame to obtain the estimated coordinates of 45 joints. Since there are n perspectives, each joint will have n estimated coordinates;
[0012] Step 5: For the estimated coordinates of 45 joints in m frames obtained from n different perspectives, use DBSCAN clustering to take the largest clustering center to form new 3D coordinates 3d_joints_new of 45 joints in m frames;
[0013] Step 6: Extract the 3D coordinates 3d_joints of 45 joints in m frames from 1_best1.pkl in Step 3. Similarly, perform global_orient processing to unify the coordinate system, and conduct error analysis with the 3d_joints_new obtained in Step 5. If the number of joints with an error exceeding 10% exceeds 20%, delete the joint spatial coordinates 3d_joints and joint rotation matrix body_pose of this frame, and use spline interpolation to supplement the deleted data to obtain 1_best2.pkl;
[0014] Step 7: Smooth 1_best2.pkl through the smoothnet smoothing network to obtain the final optimized data 1_best3.pkl;
[0015] Step 8: Process 1_best3.pkl through the SMPLX plugin and Python file in Blender to obtain the fbx action reconstruction file format 1_best3.fbx.
[0016] Furthermore, the specific method of Step 1 is as follows: When photographing the target person, use a synchronous signal generator to connect all cameras to the same synchronous signal source to ensure that all cameras start and stop recording at the same time. The cameras are arranged in a ring and evenly distributed around the target person on the circumference to ensure that the viewing angles of each camera cover different sides of the target person.
[0017] Furthermore, the specific method of Step 2 is as follows:
[0018] Step 2.1) When inputting the video frame sequence, the system performs human detection through the get_detections method, and uses the adaptive bounding box expansion algorithm to adjust the detection area to a fixed aspect ratio W:H = 3:4. The expressions for the width and height expansion amounts are:
[0019] Step 2.2) Use the HMR2023TextureSampler module that captures smpl model parameters to process the cropped image area and extract the smpl model of the video person, which includes shape parameters and pose parameters
[0020] Step 2.3) In the geometric projection stage, the system converts the predicted vertices to the image plane through perspective projection, and then synthesizes the texture atlas;
[0021] Step 2.4) In the feature extraction stage, the network projects the detection area to the UV texture space for texture mapping through the differential renderer, and establishes the mapping relationship between pixel coordinates and 3D mesh vertices v uv = K(R·v3d + t), where K is the intrinsic parameter matrix, R is a 3×3 camera rotation matrix, t is a 3×1 camera translation vector, and v 3d represents the vertex coordinates of the SMPL model, and v uv is the UV texture space coordinates of the coordinate mapping;
[0022] Step 2.5) The temporal optimization module uses a sliding window strategy to kinematically smooth the pose parameters and shape parameters between consecutive frames by minimizing the energy function to ensure parameter continuity, where θ t and β t are the pose and shape parameters of the t-th frame, is the global average shape parameter, and λ is the regularization coefficient that balances the pose parameters and shape parameters;
[0023] The final output data is encoded and compressed to generate a pkl file, and its data structure contains a pose matrix aligned with timestamps a matrix composed of shape vectors 3D joint coordinates and tracking identifiers to achieve an end-to-end conversion from the original video to a parametric representation.
[0024] Furthermore, the specific method of step 3 is as follows:
[0025] Step 3.1) Perform standardization preprocessing on the original shape parameters. Let the shape vector of each frame be an l-dimensional column vector β, and the shape parameters of m frames be represented as a matrix B. The shape vector β i after standardization is represented as:
[0026]
[0027] where μ i = E(β i ) is the mean of the shape vector, and σ i = Var(β i ) is the standard deviation. Through this operation, the standardized data B std is obtained;
[0028] Step 3.2) Determine the principal component direction by constructing a covariance matrix and solve the optimization problem:
[0029]
[0030] where w is the unit vector of the projection direction, and its solution is the eigenvector corresponding to the largest eigenvalue of C. The first k eigenvectors are selected in descending order of eigenvalues to form an orthogonal basis matrix V k , and the value of k is determined by the cumulative variance contribution rate Determined to reach the preset threshold, λ i The k-th eigenvalue reflects the variance intensity of the data in the corresponding direction;
[0031] Step 3.3) Achieve dimensionality reduction and reconstruction through projection. Multiply the standardized data B std on the right by the orthogonal basis matrix V k to obtain the dimension-reduced data B pca , and then multiply on the right by to achieve back-projection reconstruction to obtain B recon , and finally perform inverse standardization to restore the data scale to obtain the denoised shape vector:
[0032]
[0033] where is the shape vector of the i-th frame after reconstruction.
[0034] Furthermore, the specific method of step 4 is as follows:
[0035] Step 4.1) Structurally describe the multi-view joint data. Suppose videos of the same action are collected from n views, and each video contains m frames. Extract the 3D joint data of each view through the SMPL model. Among them, the 3D joint position matrix of the i-th view contains the three-dimensional coordinates of m frames and 45 joints. Perform spatial alignment through the global rotation matrix provided by the SMPL model. R i,t represents the global rotation matrix from the reference coordinate system to the local coordinate system of the t-th frame of the i-th view, which is estimated by the SMPL model according to the sensor pose;
[0036] Step 4.2) Inversely map the local coordinates back to the reference coordinate system through coordinate transformation. The coordinate transformation formula is:
[0037]
[0038] where represents the local coordinates of the j-th joint of the i-th view and the t-th frame. J ′ i,j (t) is the global coordinate in the reference coordinate system;
[0039] Step 4.3) After independently transforming each frame of data from all views, obtain the global joint position set J ′ i (t) = {J ′ i,j (t)|1 ≤ j ≤ 45}, 1 ≤ i ≤ n, 1 ≤ t ≤ m.
[0040] Further, the specific method of step 5 is as follows:
[0041] Step 5.1) For the estimated coordinates J′1(t), J′2(t)…J′ n (t) of n perspectives of all joints in the t-th frame, use the DBSCAN algorithm for clustering, find the largest cluster in the t-th frame and calculate its center to obtain where is the cluster center coordinate of the j-th joint in the t-th frame, t = 1, 2, …, m;
[0042] Step 5.2) Store the cluster centers obtained from analyzing 45 joints in a total of m frames in a dictionary, create an empty 3×45×m tensor 3d_joints_new, and fill into the corresponding positions of 3d_joints_new. The final elements of 3d_joints_new are expressed as where t represents the frame number, t = 1, 2, …, m.
[0043] Further, the specific method of step 6 is as follows:
[0044] Step 6.1) Extract the 45 three-dimensional joint point coordinates of the t-th frame from the initial data file 1_best1.pkl Perform coordinate system normalization processing through the global rotation matrix of the t-th frame, and map the joint points to the world coordinate system:
[0045]
[0046] where J t is the original joint point coordinate, R t is the global rotation matrix, ensuring strict alignment of the data with 3d_joints_new in the Euclidean space. Calculate the frame-by-frame joint residual matrix and quantify the error through the Frobenius norm:
[0047]
[0048] where e t,ij is the element of the residual matrix E t and represents the error of the i-th joint in the t-th frame on the j-th coordinate axis;
[0049] Step 6.2) Introduce the relative error index where the denominator term characterizes the overall scale information of the optimized joint point set, and identify abnormal frames by setting a threshold function:
[0050]
[0051] The determination condition is satisfied That is, when more than 20% of the joint relative error indicators in a single frame exceed 10%, the culling mechanism is triggered to delete the joint coordinates 3d_joints and the joint rotation matrix body_pose of this frame;
[0052] Step 6.3) For the missing frame joint coordinates 3d_joints and the joint rotation matrix body_pose, use cubic spline interpolation to construct a piecewise polynomial function S(x) to ensure that the interpolation result is smooth and continuous;
[0053] Step 6.4) Supplement the interpolated joint coordinates and rotation matrix into the data after deletion to generate complete 3d_joints and body_pose data, and save them as the 1_best2.pkl file.
[0054] Furthermore, the specific method of step 7 is as follows:
[0055] Step 7.1) After obtaining the initial action parameter file 1_best2.pkl, first convert the rotation matrix representation of the SMPL model to the axis-angle representation to form a sequence of rotation vectors Adopt an interleaved frame processing strategy: decompose the sequence of rotation vectors Ω into an odd frame subset Ω odd ={ω 2k+1} and an even frame subset Ω even ={ω 2k}, and input them into the SmoothNet smoothing network for independent optimization to obtain and
[0056] Step 7.2) Decompose the three-dimensional translation vector t t =(x t , y t , z t ) extracted from the pkl file into horizontal displacement (x t , y t ) and vertical displacement z t for independent processing, that is, minimize the height constraint energy function where z t is the vertical displacement of the t-th frame, is the global average vertical displacement, and λ is the regularization coefficient used to control the smoothness degree. Keep reasonable contact between the human body and the ground during the smoothing process, and then reconstruct the complete time series through the interleaved recombination algorithm for the optimized odd and even subsequences and ;
[0057] Step 7.3) Inverse-transform the axis-angle representation into a rotation matrix through the Rodriguez formula, reconstruct the pose parameters conforming to the SMPL model topology, and output the optimized motion parameter file 1_best3.pkl.
[0058] Furthermore, the specific method of Step 8 is as follows:
[0059] Step 8.1) Read the motion parameter file 1_best3.pkl in Blender to obtain the human body pose parameters The smpl model parameters of the shape parameters, forming a time series parameter set {θ t , β t , γ t}, where t = 1, …, T, and and respectively represent the smpl pose parameters, shape parameters, and root joint translation of the t-th frame;
[0060] Step 8.2) Call the Python code provided by the 4dhuman project for generating fbx motion files through the Blender SMPLX plugin, process the extracted smpl model parameters to obtain the skinned mesh, bone hierarchy, and time series motion information, and write the above information into 1_best3.fbx to support direct driving of character animations by the Unity / Unreal engine.
[0061] Beneficial effects: Compared with the prior art, the technical solution of the present invention has the following beneficial effects:
[0062] The present invention significantly improves the accuracy and stability of the 4DHuman motion capture technology through the method of multi-view video acquisition combined with post-processing network optimization. In the process of multi-view data fusion, the present invention uses clustering analysis to optimize the joint data, effectively solving the problem of multi-view data consistency. Especially when the number of views is large, the clustering center of the joint is accurately extracted through the DBSCAN clustering algorithm, ensuring the accuracy and reliability of the data. In addition, the motion capture results are further optimized by error analysis and spline interpolation to supplement the data, avoiding motion distortion caused by data noise or outliers.
[0063] The present invention introduces the SmoothNet smoothing network to process the motion data, significantly improving the smoothness of the motion. SmoothNet smooths the time series data through a sliding window strategy, effectively removing the jitter and noise in the motion parameters, making the human body motion more natural and coherent. This method not only improves the quality of the motion capture data but also provides high-quality input data for subsequent animation production and virtual reality applications.
[0064] The present invention realizes the reconstruction of human actions by inputting the optimized data file into Blender and using the SMPLX plug-in, and further exports the reconstruction result as an FBX file and applies it to the Unity platform. This process not only realizes the efficient conversion from multi-view videos to high-quality 3D actions, but also provides powerful technical support for fields such as virtual reality, animation production, and game development. Compared with the prior art, the present invention has significant advantages in terms of data accuracy, action smoothness, and application compatibility, and has broad application prospects and practical application value. Description of the Drawings
[0065] Figure 1 It is a specific flowchart of the method of the present invention. Detailed Embodiments
[0066] In order to make the objectives, content, and advantages of the present invention clearer, the details in the present invention will be further described below with reference to the drawings and examples.
[0067] As Figure 1 shown, the present invention proposes a post-processing action reconstruction method for 4D human motion capture technology based on multi-view video acquisition. The method includes the following steps:
[0068] Step 1, photograph the target person from n perspectives to obtain multiple perspective video clips of the target person, where n≥2, and take one of them as the target perspective and the others as auxiliary perspectives;
[0069] Step 2, input the above n video clips into the 4D Human human motion estimation network to obtain the corresponding n pkl files;
[0070] Step 3, extract the shape parameters betas of the human models in a total of m frames in the pkl file of the target perspective, form a sample matrix with m rows, and perform noise reduction processing to obtain the optimized pkl file 1_best1.pkl;
[0071] Step 4, extract the 3D coordinates 3d_joints of 45 joints in m frames from the n pkl files, and unify the coordinate system after being processed by the global rotation matrix globle_orient of this frame to obtain the estimated coordinates of 45 joints. Since there are n perspectives, each joint will have n estimated coordinates;
[0072] Step 5, for the estimated coordinates of 45 joints in m frames obtained from n different perspectives, use DBSCAN clustering to take the largest clustering center to form new 3D coordinates 3d_joints_new of 45 joints in m frames;
[0073] Step 6: Extract the 3D coordinates of 45 joints in m frames in 1_best1.pkl in Step 3. Similarly, perform global_orient processing to unify the coordinate system, and conduct error analysis with the 3d_joints_new obtained in Step 5. If the number of joints with an error exceeding 10% exceeds 20%, delete the joint spatial coordinates 3d_joints and the joint rotation matrix body_pose of this frame, and use spline interpolation to supplement the deleted data to obtain 1_best2.pkl;
[0074] Step 7: Smooth 1_best2.pkl through the smoothnet smoothing network to obtain the final optimized data 1_best3.pkl;
[0075] Step 8: Process 1_best3.pkl through the SMPLX plugin and python file in blender to obtain the fbx action reconstruction file format 1_best3.fbx.
[0076] Furthermore, the specific method of Step 1 is as follows: When photographing the target person, use a synchronous signal generator to connect all cameras to the same synchronous signal source to ensure that all cameras start and stop recording at the same time. The cameras are arranged in a ring and evenly distributed around the target person on the circumference to ensure that the viewing angles of each camera cover different sides of the target person.
[0077] Furthermore, the specific method of Step 2 is as follows:
[0078] Step 2.1): When inputting the video frame sequence, the system performs human detection through the get_detections method, and uses the adaptive bounding box expansion algorithm to adjust the detection area to a fixed aspect ratio of W:H = 3:4. The expressions for the width and height expansion amounts are:
[0079] Step 2.2): Use the HMR2023TextureSampler module that captures smpl model parameters to process the cropped image area and extract the smpl model of the video person, including the shape parameters and pose parameters
[0080] Step 2.3): In the geometric projection stage, the system converts the predicted vertices to the image plane through perspective projection, and then synthesizes the texture atlas;
[0081] Step 2.4): In the feature extraction stage, the network projects the detection area to the UV texture space for texture mapping through the differential renderer, and establishes the mapping relationship v between the pixel coordinates and the 3D mesh vertices uv = K(R·v3d + t), where K is the intrinsic parameter matrix, R is a 3×3 camera rotation matrix, t is a 3×1 camera translation vector, and v 3d represents the vertex coordinates of the SMPL model, and v uv are the UV texture space coordinates of the coordinate mapping;
[0082] Step 2.5) The temporal optimization module uses a sliding window strategy to perform kinematic smoothing on the pose parameters and shape parameters between consecutive frames, and minimizes the energy function to ensure parameter continuity, where θ t and β t are the pose and shape parameters of the t-th frame, is the global average shape parameter, and λ is the regularization coefficient that balances the pose parameters and shape parameters;
[0083] The final output data is encoded and compressed to generate a pkl file, and its data structure contains the pose matrix aligned with the timestamp The matrix composed of shape vectors 3D joint coordinates and the tracking identifier to achieve an end-to-end conversion from the original video to the parametric representation.
[0084] Furthermore, the specific method of step 3 is as follows:
[0085] Step 3.1) Perform standardization preprocessing on the original shape parameters. Let the shape vector of each frame be an l-dimensional column vector β, and the shape parameters of m frames be represented as a matrix B. The shape vector β i after standardization is expressed as:
[0086]
[0087] where μ i = E(β i ) is the mean of the shape vector, and σ i = Var(β i ) is the standard deviation. Through this operation, the standardized data B std is obtained;
[0088] Step 3.2) Determine the principal component direction by constructing a covariance matrix and solve the optimization problem:
[0089]
[0090] where w is the unit vector of the projection direction, and its solution is the eigenvector corresponding to the largest eigenvalue of C. The first k eigenvectors are selected in descending order of eigenvalues to form an orthogonal basis matrix V k , and the value of k is determined by the cumulative variance contribution rate Determined to reach the preset threshold, λ i The k-th eigenvalue reflects the variance intensity of the data in the corresponding direction;
[0091] Step 3.3) Achieve dimensionality reduction and reconstruction through projection. Multiply the standardized data B std on the right by the orthogonal basis matrix V k to obtain the dimensionality-reduced data B pca , and then multiply on the right by to achieve back-projection reconstruction to obtain B recon , and finally perform anti-standardization to restore the data scale to obtain the denoised shape vector:
[0092]
[0093] where is the shape vector of the i-th frame after reconstruction.
[0094] Furthermore, the specific method of step 4 is as follows:
[0095] Step 4.1) Structurally describe the multi-view joint data. Assume that videos of the same action are collected from n perspectives, and each video contains m frames. Extract the 3D joint data of each perspective through the SMPL model. Among them, the 3D joint position matrix of the i-th perspective contains the three-dimensional coordinates of m frames and 45 joints. Align the space through the global rotation matrix provided by the SMPL model. R i,t represents the global rotation matrix from the reference coordinate system to the local coordinate system of the t-th frame of the i-th perspective, which is estimated by the SMPL model according to the sensor pose;
[0096] Step 4.2) Inversely map the local coordinates back to the reference coordinate system through coordinate transformation. The coordinate transformation formula is:
[0097]
[0098] where represents the local coordinates of the j-th joint in the i-th perspective and the t-th frame, and J′ i,j (t) is the global coordinate in the reference coordinate system;
[0099] Step 4.3) After independently transforming each frame of data from all perspectives, obtain the global joint position set J′ i (t) = {J′ i,j (t)|1 ≤ j ≤ 45}, 1 ≤ i ≤ n, 1 ≤ t ≤ m.
[0100] Furthermore, the specific method of step 5 is as follows:
[0101] Step 5.1) For the estimated coordinates J′1(t), J′2(t).....J′ n (t) of all joints in the t-th frame from n perspectives, use the DBSCAN algorithm for clustering, find the largest cluster in the t-th frame and calculate its center to obtain where, is the cluster center coordinate of the j-th joint in the t-th frame, t = 1, 2, …, m;
[0102] Step 5.2) Store the cluster centers obtained from analyzing 45 joints in a total of m frames in a dictionary, create an empty 3×45×m tensor 3d_joints_new, and fill into the corresponding positions of 3d_joints_new. The final elements of 3d_joints_new are expressed as where, t represents the frame number, t = 1, 2, …, m.
[0103] Furthermore, the specific method of Step 6 is as follows:
[0104] Step 6.1) Extract the 45 three-dimensional joint point coordinates of the t-th frame from the initial data file 1_best1.pkl Perform coordinate system normalization processing through the global rotation matrix of the t-th frame to map the joint points to the world coordinate system:
[0105]
[0106] where, J t is the original joint point coordinate, R t is the global rotation matrix, ensure that the data is strictly aligned with 3d_joints_new in the Euclidean space, calculate the frame-by-frame joint residual matrix and quantify the error through the Frobenius norm:
[0107]
[0108] where, e t,ij is the element of the residual matrix E t and represents the error of the i-th joint in the t-th frame on the j-th coordinate axis;
[0109] Step 6.2) Introduce the relative error index where the denominator term characterizes the overall scale information of the optimized joint point set, and identify abnormal frames by setting a threshold function:
[0110]
[0111] Its judgment condition is satisfied That is, when more than 20% of the joint relative error indicators in a single frame exceed 10%, the culling mechanism is triggered to delete the joint coordinates 3d_joints and the joint rotation matrix body_pose of this frame;
[0112] Step 6.3) For the missing frame's joint coordinates 3d_joints and joint rotation matrix body_pose, use cubic spline interpolation to construct a piecewise polynomial function S(x) to ensure the interpolation result is smooth and continuous;
[0113] Step 6.4) Supplement the interpolated joint coordinates and rotation matrix into the data after deletion to generate complete 3d_joints and body_pose data, and save them as the 1_best2.pkl file.
[0114] Furthermore, the specific method of step 7 is as follows:
[0115] Step 7.1) After obtaining the initial action parameter file 1_best2.pkl, first convert the rotation matrix representation of the SMPL model to the axis-angle representation to form a sequence of rotation vectors Adopt an interleaved frame processing strategy: decompose the sequence of rotation vectors Ω into an odd-frame subset Ω odd ={ω 2k+1} and an even-frame subset Ω even ={ω 2k}, and input them into the SmoothNet smoothing network for independent optimization to obtain and
[0116] Step 7.2) Decompose the three-dimensional translation vector t t =(x t ,y t ,z t ) extracted from the pkl file into the horizontal displacement (x t ,y t ) and the vertical displacement z t for independent processing, that is, minimize the height constraint energy function where z t is the vertical displacement of the t-th frame, is the global average vertical displacement, and λ is the regularization coefficient used to control the smoothness. During the smoothing process, keep the reasonable contact between the human body and the ground, and then reconstruct the complete time series through the interleaved recombination algorithm for the optimized odd and even subsequences and ;
[0117] Step 7.3) Inverse-transform the axis-angle representation into a rotation matrix through the Rodriguez formula, reconstruct the pose parameters that conform to the SMPL model topology, and output the optimized motion parameter file 1_best3.pkl.
[0118] Furthermore, the specific method of Step 8 is as follows:
[0119] Step 8.1) Read the motion parameter file 1_best3.pkl in Blender to obtain the human pose parameters The SMPL model parameters of the shape parameters to form a time series parameter set {θ t , β t , γ t}, where t = 1, …, T, and and respectively represent the SMPL pose parameters, shape parameters, and root joint translation of the t-th frame;
[0120] Step 8.2) Call the Python code provided by the 4dhuman project for generating fbx motion files through the Blender SMPLX plugin, process the extracted SMPL model parameters to obtain the skinned mesh, bone hierarchy, and time series motion information, and write the above information into 1_best3.fbx to support direct driving of character animations by the Unity / Unreal engine.
Claims
1. A post - processing action reconstruction method for 4D human motion capture technology based on multi - perspective video acquisition, characterized in that, The method includes the following steps: Step 1: Shoot the target person from n perspectives to obtain multiple perspective video clips of the target person, where n≥2. One of them is used as the target perspective, and the others are used as auxiliary perspectives; Step 2: Input the above n video clips into the 4DHuman human motion estimation network to obtain the corresponding n pkl files; Step 3: Extract the shape parameters betas of the human model of a total of m frames in the pkl file of the target perspective to form a sample matrix with m rows, and perform noise reduction processing to obtain the optimized pkl file 1_best1.pkl; Step 4: Extract the 3D coordinates 3d_joints of 45 joints in m frames from the n pkl files, and unify the coordinate system after being processed by the global rotation matrix globle_orient of this frame to obtain the estimated coordinates of 45 joints. Since there are n perspectives, each joint will have n estimated coordinates; Step 5: For the estimated coordinates obtained from n different perspectives of 45 joints in m frames, use DBSCAN clustering to take the largest clustering center to form new 3D coordinates 3d_joints_new of 45 joints in m frames; Step 6: Extract the 3D coordinates 3d_joints of 45 joints in m frames in 1_best1.pkl in Step 3. Similarly, unify the coordinate system after being processed by globle_orient, and perform error analysis on it and 3d_joints_new obtained in Step 5. If the number of joints with an error exceeding 10% exceeds 20%, then delete the joint space coordinates 3d_joints and the joint rotation matrix body_pose of this frame, and use spline interpolation to supplement the deleted data to obtain 1_best2.pkl; Step 7: Smoothly process 1_best2.pkl through the smoothnet smoothing network to obtain the final optimized data 1_best3.pkl; Step 8: Process 1_best3.pkl through the SMPLX plug-in and python file in blender to obtain the fbx motion reconstruction file format 1_best3.fbx.
2. The post - processing network optimization method for a 4D human motion capture technology based on multi - perspective video acquisition according to claim 1, characterized in that, The specific method of Step 1 is as follows: When shooting the target person, use a synchronous signal generator to connect all cameras to the same synchronous signal source to ensure that all cameras start and stop recording at the same time. The cameras are arranged in a ring and evenly distributed on the circumference around the target person to ensure that the perspective of each camera covers different sides of the target person.
3. The post - processing network optimization method for a 4D human motion capture technology based on multi - perspective video acquisition according to claim 2, wherein, The specific method of Step 2 is as follows: Step 2.1) When inputting a video frame sequence, the system performs human detection through the get_detections method and adjusts the detection area to a fixed aspect ratio of W:H = 3:4 using an adaptive bounding box expansion algorithm. The expressions for the width and height expansion amounts are as follows: Step 2.2) Use the HMR2023TextureSampler module that captures the SMPL model parameters to process the cropped image region and extract the SMPL model of the video person, where the shape parameters are included and pose parameters Step 2.3) In the geometric projection stage, the system converts the predicted vertices to the image plane through perspective projection, and then synthesizes the texture map; Step 2.4) In the feature extraction stage, the network projects the detection region onto the UV texture space for texture mapping through a differential renderer, and establishes the mapping relationship v between pixel coordinates and 3D mesh vertices uv = K(R · v 3d + t), where K is the intrinsic parameter matrix, R is a 3×3 camera rotation matrix, t is a 3×1 camera translation vector, and v 3d represents the SMPL model vertex coordinates, and v uv is the UV texture space coordinate of this coordinate mapping; Step 2.5) The timing optimization module uses a sliding window strategy to perform kinematic smoothing on the pose parameters and shape parameters between consecutive frames, by minimizing the energy function to ensure parameter continuity, where θ t and β t are the pose and shape parameters of the t-th frame, is the global average shape parameter, and λ is the regularization coefficient that balances the pose parameters and shape parameters; The final output data is encoded and compressed to generate a pkl file, and its data structure contains pose matrices aligned with timestamps A matrix composed of shape vectors 3D joint coordinates and tracking identifiers To achieve an end-to-end conversion from the original video to a parametric representation.
4. The post - processing network optimization method for a 4D human motion capture technology based on multi - perspective video acquisition according to claim 3, characterized in that, The specific method of Step 3 is as follows: Step 3.1) Perform standardized preprocessing on the original shape parameters. Let the shape vector of each frame be an l-dimensional column vector β, and the shape parameters of m frames be represented as a matrix B. The shape vector β of the i-th frame i After standardization, it is expressed as: Among them, μ i = E(β i ) is the mean of the shape vectors, and σ i = Var(β i ) is the standard deviation. After this operation, the standardized data B std ; Step 3.2) By constructing a covariance matrix Determine the principal component direction and solve the optimization problem: Among them, w is the unit vector of the projection direction, and its solution is the eigenvector corresponding to the largest eigenvalue of C. The first k eigenvectors are selected in descending order of eigenvalues to form an orthogonal basis matrix V k , and the value of k is determined by the cumulative variance contribution rate reaching a preset threshold. λ i is the k-th eigenvalue, which reflects the variance intensity of the data in the corresponding direction; Step 3.3) Achieve dimensionality reduction and reconstruction through projection, and for the standardized data B std right multiply by the orthogonal basis matrix V k to obtain the dimensionally reduced data B pca , then right multiply by to achieve back-projection reconstruction to obtain B recon , and finally perform inverse standardization to restore the data scale to obtain the denoised shape vector: Among them, is the shape vector of the i-th frame after reconstruction.
5. The post - processing network optimization method for a 4D human motion capture technology based on multi - perspective video acquisition according to claim 4, characterized in that, The specific method of Step 4 is as follows: Step 4.1) Structurally describe the multi-view joint data. Assume that videos of the same action are collected from n views, and each video contains m frames. Extract the 3D joint data of each view through the SMPL model. Among them, the 3d_joints joint position matrix of the i-th view contains the three-dimensional coordinates of m frames and 45 joints, and performs spatial alignment through the global rotation matrix provided by the SMPL model. R i,t represents the global rotation matrix from the reference coordinate system to the local coordinate system of the t-th frame of the i-th view, which is estimated by the SMPL model according to the sensor pose; Step 4.2) Inverse map the local coordinates back to the reference coordinate system through coordinate transformation. The coordinate transformation formula is: Among them, represents the local coordinates of the j-th joint in the i-th perspective and the t-th frame, and J′ i,j (t) is the global coordinate in the reference coordinate system; Step 4.3) After independently transforming each frame of data from all viewpoints, a global joint position set J′ in a unified coordinate system is obtained. i J′ i,j (t) = {J′ (t)|1 ≤ j ≤ 45}, 1 ≤ i ≤ n, 1 ≤ t ≤ m.
6. The post - processing network optimization method for a 4D human motion capture technology based on multi - perspective video acquisition according to claim 5, characterized in that, The specific method of Step 5 is as follows: Step 5.1) For the estimated coordinates \(J'_1(t), J'_2(t), \cdots, J'_{n}(t)\) of all joints in the \(t\)-th frame from \(n\) viewpoints, use the DBSCAN algorithm for clustering, find the largest cluster in the \(t\)-th frame and calculate its center to obtain n (t), where the largest cluster in the \(t\)-th frame is found by clustering using the DBSCAN algorithm and its center is calculated to obtain where is the cluster center coordinate of the \(j\)-th joint in the \(t\)-th frame, \(t = 1, 2, \cdots, m\); Step 5.2) Store the cluster centers obtained from analyzing 45 joints in a total of m frames in a dictionary, create an empty 3×45×m tensor 3d_joints_new, and fill it into the corresponding positions of 3d_joints_new. The elements of the finally obtained 3d_joints_new are expressed as where t represents the frame number, t = 1, 2, …, m.
7. The post - processing network optimization method for a 4D human motion capture technology based on multi - perspective video acquisition according to claim 6, characterized in that, The specific method of Step 6 is as follows: Step 6.1) Extract the 45 three-dimensional joint point coordinates of the t-th frame from the initial data file 1_best1.pkl Perform coordinate system normalization through the global rotation matrix of the t-th frame and map the joint points to the world coordinate system: Among them, J t is the original joint point coordinate, and R t is the global rotation matrix, which ensures that the data is strictly aligned with 3d_joints_new in the Euclidean space. Calculate the frame-by-frame joint residual matrix and quantify the error through the Frobenius norm: Among them, e t,ij is an element of the residual matrix E t , representing the error of the i-th joint in the j-th coordinate axis at the t-th frame; Step 6.2) Introduce the relative error index where the denominator term represents the overall scale information of the optimized joint point set, and abnormal frames are identified by setting a threshold function: The determination condition is satisfied That is, when more than 20% of the joint relative error indicators within a single frame exceed 10%, the culling mechanism is triggered to delete the joint coordinates 3d_joints and the joint rotation matrix body_pose of this frame; Step 6.3) For the joint coordinates 3d_joints and joint rotation matrices body_pose of the missing frames, use cubic spline interpolation to construct a piecewise polynomial function S(x) to ensure that the interpolation results are smooth and continuous. Step 6.4) Supplement the interpolated joint coordinates and rotation matrices into the data after deletion to generate complete 3d_joints and body_pose data, and save them as the 1_best2.pkl file.
8. The post - processing network optimization method for a 4D human motion capture technology based on multi - perspective video acquisition according to claim 7, characterized in that, The specific method of Step 7 is as follows: Step 7.1) After obtaining the initial action parameter file 1_best2.pkl, first represent the rotation matrix of the SMPL model is converted to the axis-angle representation to form a sequence of rotation vectors Adopt an interleaved frame processing strategy: decompose the rotation vector sequence Ω into an odd-frame subset Ω odd ={ω 2k+1} and an even-frame subset Ω even ={ω 2k}, and input them into the SmoothNet smoothing network for independent optimization to obtain and Step 7.2) Decompose the three-dimensional translation vector t t =(x t , y t , z t ) extracted from the pkl file into horizontal displacement (x t , y t ) and vertical displacement z t for independent processing, that is, minimize the height constraint energy function where z t is the vertical displacement of the t-th frame, is the global average vertical displacement, λ is the regularization coefficient used to control the smoothness, and maintain reasonable contact between the human body and the ground during smoothing. Then, reconstruct the complete time series by interleaving and recombining the optimized odd and even subsequences and through the interleaving recombination algorithm; Step 7.3) Inverse transform the axis-angle representation into a rotation matrix through the Rodriguez formula, reconstruct the pose parameters that conform to the SMPL model topology, and output the optimized motion parameter file 1_best3.pkl.
9. The post - processing network optimization method for a 4D human motion capture technology based on multi - perspective video acquisition according to claim 8, characterized in that, The specific method of Step 8 is as follows: Step 8.1) Read the motion parameter file 1_best3.pkl in Blender to obtain the human body pose parameters The SMPL model parameters of the shape parameters to form a time series parameter set {θ t , β t , γ t}, where t = 1, …, T, and and represent the SMPL pose parameters, shape parameters, and root joint translation of the t-th frame, respectively. Step 8.2) Call the Python code provided by the 4dhuman project for generating fbx motion files through the Blender SMPLX plugin, process the extracted smpl model parameters to obtain the skinned mesh, bone hierarchy, and time series motion information, and write the above information into 1_best3.fbx to support direct driving of character animations by the Unity / Unreal engines.