A multi-view non-calibration based human motion fast capturing method and system

By employing a multi-view calibration-free method and a physical geometry consistency strategy, and utilizing neural networks for automatic calibration and iterative calculation, the problems of cumbersome calibration and insufficient accuracy in multi-view human motion capture are solved, achieving efficient, real-time, and high-precision human motion capture.

CN120852679BActive Publication Date: 2026-02-03HUNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511349651.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-22
Publication Date
2026-02-03
Estimated Expiration
2045-09-22

AI Technical Summary

Technical Problem

Existing human motion capture technology suffers from problems such as cumbersome calibration, high hardware dependence, poor adaptability to dynamic environments, and insufficient reconstruction accuracy in multi-view reconstruction, making it difficult to meet the low latency and high precision requirements of fields such as virtual reality and smart healthcare.

Method used

A multi-view calibration-free method based on neural networks is adopted to achieve automatic calibration through multi-view video frames. The camera extrinsic parameters are calculated iteratively using a physical geometry consistency strategy, and 3D reconstruction is performed by combining a skinned multi-person linear model, outputting high-precision digital 3D human motion.

Benefits of technology

It achieves efficient human motion capture without cumbersome calibration work, combining real-time performance and high precision, and is suitable for small and medium-sized studios and consumer-grade scenarios, improving the robustness and computational efficiency of multi-view motion capture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120852679B_ABST
    Figure CN120852679B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on multi-view uncalibrated human action fast capture method and system, method includes: using the human three-dimensional reconstruction model based on neural network and training, according to the video of each different view, the human three-dimensional reconstruction of corresponding view is carried out, obtains the digital three-dimensional human action under each camera coordinate system;Initial external parameter is estimated to each different view camera, then the initial external parameter of each camera is used and the digital three-dimensional human action under corresponding camera coordinate system, and based on physical geometry consistency strategy, each camera external parameter is iteratively calculated;The camera external parameter obtained by iterative calculation is used, the digital three-dimensional human action under corresponding camera coordinate system is converted to world coordinate system, and finally the digital three-dimensional human action is obtained by multi-view fitting.The application does not need heavy calibration work, realizes automatic calibration using multi-view video frame, and then quickly captures human action and completes three-dimensional reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing, specifically to a method and system for rapid human motion capture based on multi-view uncalibrated images. Background Technology

[0002] Human motion capture technology has wide-ranging applications in virtual reality, film and television special effects, motion analysis, and medical rehabilitation. Traditional motion capture methods mainly rely on optical markers, inertial sensors, or depth cameras. While these methods offer high accuracy, they typically require complex calibration processes, expensive hardware, or wearable sensors, limiting their application in natural interactions and low-cost scenarios. In recent years, vision-based markerless motion capture technology has gained widespread attention due to its non-invasiveness and low cost, with single-view 3D reconstruction, camera calibration, and multi-view fusion being key technological directions. However, existing methods still face many challenges, necessitating a more efficient and robust markerless multi-view motion capture solution.

[0003] In single-view human 3D reconstruction, early research mainly relied on parametric human models (such as SMPL) and 2D keypoint detection to recover 3D pose and shape. However, single-view methods suffer from depth ambiguity and self-occlusion issues, limiting reconstruction accuracy under complex movements. In recent years, the introduction of deep learning techniques has significantly improved the performance of single-view reconstruction, such as the temporally optimized OpenPose improvement method, hierarchical Gaussian modeling (MultiGO), and geometric affine fields (Niagara framework). However, these methods still have shortcomings in spatiotemporal consistency modeling in dynamic scenes, making it difficult to meet the requirements of high-precision motion capture.

[0004] Camera calibration is fundamental to multi-view systems. Traditional methods rely on checkerboard patterns or calibration objects for precise parameter calculations, but these are cumbersome and difficult to adapt to dynamic environments. Self-calibration techniques attempt to estimate camera parameters using scene features (such as vanishing points), but their robustness is poor and they are easily affected by lighting and occlusion. In recent years, zero-loss machine calibration and timestamp-synchronized multi-camera acquisition methods have simplified the calibration process to some extent, but they still cannot completely eliminate manual intervention. The core challenge of multi-view systems lies in how to achieve geometrically consistent alignment of cross-view data without relying on precise calibration, especially in dynamic interactive scenarios where the accumulation of calibration errors can lead to reconstruction failure.

[0005] Multi-view technology, through collaborative observation using multiple cameras, can effectively alleviate the occlusion and ambiguity problems of single-view methods. Firstly, the calibration-free approach reduces hardware dependence, making motion capture technology easier to deploy in small and medium-sized studios and consumer-grade scenarios. Secondly, multi-view fusion overcomes the physical limitations of single-view methods, providing more robust reconstruction results in occluded and complex interactive scenarios. Finally, real-time and efficient algorithm design can meet the dual requirements of low latency and high accuracy in applications such as virtual reality and smart healthcare. In recent years, multi-view calibration-free technology has shown potential in biomechanical analysis and virtual character-driven scenarios, but systematic optimization is still needed to balance computational efficiency and reconstruction accuracy. Therefore, researching fast human motion capture methods based on multi-view calibration-free methods not only has significant academic value but will also promote innovative development in fields such as film production, smart healthcare, and human-computer interaction. Summary of the Invention

[0006] This invention provides a method and system for rapid human motion capture based on multi-view calibration-free technology. It eliminates the need for cumbersome calibration work, utilizes multi-view video frames to achieve automatic calibration, and then rapidly captures human motion and completes 3D reconstruction.

[0007] To achieve the above technical objectives, the present invention adopts the following technical solution:

[0008] A method for rapid human motion capture based on multi-view, uncalibrated motion capture includes:

[0009] Using a neural network-based and trained 3D human body reconstruction model, 3D human body reconstruction is performed from the perspective of the video from different angles to obtain digital 3D human body motion in each camera coordinate system.

[0010] An initial extrinsic parameter is estimated for each camera from a different perspective. Then, using the initial extrinsic parameter of each camera and the digital 3D human motion in the corresponding camera coordinate system, and based on a physical-geometric consistency strategy, the extrinsic parameter of each camera is iteratively calculated.

[0011] Using the camera extrinsic parameters obtained through iterative calculation, the digital 3D human motion in the corresponding camera coordinate system is transformed to the world coordinate system, and then the final digital 3D human motion is obtained through multi-view fitting.

[0012] Furthermore, the digital 3D human body adopts a skinned multi-person linear model, whose parameters include the shape vector, posture vector, and offset vector of several skeletal joints of the human body.

[0013] Furthermore, the human body 3D reconstruction model obtained based on and trained by the neural network includes:

[0014] The feature extraction network employs YOLO's multi-scale feature fusion mechanism and attention enhancement to extract features from the input video frames;

[0015] Multilayer perceptron is used to reconstruct coarse-grained digital 3D human bodies based on extracted features.

[0016] The reprojection module is used to map the digital 3D human body onto the extracted features to obtain a feedback vector;

[0017] The update module, including a gated neural network and a regressive multilayer perceptron, is used to unwrap and regress the feedback vector to obtain a fine-grained digital 3D human body.

[0018] Furthermore, the feature extraction network includes multiple feature extraction layers and hierarchical feature fusion paths;

[0019] The multi-feature extraction layer adopts a three-stage hierarchical structure and constructs a multi-scale feature space through progressive downsampling: (1) shallow feature extraction, including two 3×3 standard convolutions and three C3k2 modules, outputting a 512-dimensional feature map; (2) mid-level feature extraction, obtaining 1024-dimensional features through two downsamplings, and using a C3k2 module with residual connections to enhance feature reuse efficiency; (3) deep semantic extraction, introducing an SPPF spatial pyramid pooling module to expand the receptive field, and cooperating with a dual-channel C2PSA attention module to enhance key point response features; finally, a 1024-dimensional feature map containing high-dimensional semantic information is output.

[0020] Furthermore, the loss functions used to train the human 3D reconstruction model include: 3D joint loss, 2D projection joint loss, and SMPL model update loss; among which;

[0021] The three-dimensional joint loss is denoted as The calculation formula is:

[0022] ;

[0023] in, For the true values ​​of the three-dimensional joints, These are 3D joints of a digital 3D human body, representing both coarse-grained and fine-grained types. It is the number of joints;

[0024] The two-dimensional projection joint loss is denoted as The calculation formula is:

[0025] ;

[0026] In the formula, This represents the mapping of the true values ​​of 3D joints to the current viewpoint. These are 2D joints of a digital 3D human body with coarse and fine granularity, respectively, viewed from the current perspective.

[0027] The SMPL model update loss is denoted as... The calculation formula is as follows:

[0028] ;

[0029] In the formula, and These are digital 3D human figures with coarse and fine granularity, respectively.

[0030] Furthermore, the estimation of an initial extrinsic parameter for each camera from a different viewpoint includes:

[0031] First, using video frames from various viewpoints, existing 2D pose estimation and tracking methods are employed to estimate the 2D pose of the human body and its corresponding confidence level in each viewpoint; then, the joints are... From a perspective The 2D pose and confidence obtained from the video frames are denoted as follows: , ;

[0032] Then, using the epipolar geometry of the known intrinsic parameters of the camera, the extrinsic parameters of each camera are solved from the multi-view 2D human pose and used as the initial extrinsic parameters.

[0033] Furthermore, the iterative calculation of the extrinsic parameters of each camera includes:

[0034] Using the initial extrinsic parameters of each camera, the digital 3D human body in the corresponding camera coordinate system is transformed to the world coordinate system; the joints are... From a perspective The world coordinates obtained from the video frames are as follows ;

[0035] Based on the distance between the position of the key point in the previous frame and the light ray in the current frame from the same viewpoint, the physical consistency constraint is constructed as follows:

[0036] ;

[0037] By forcing the rays from the joints to coexist in different viewpoints, a geometric consistency constraint is constructed as follows:

[0038] ;

[0039] In the formula, Representative key points From the perspective The light in the middle, It is a direction vector. This is the moment vector of the light ray, which is used to encode the position information of the light ray in space; Key points From the perspective The world coordinates relative to the previous frame t-1 in the current frame; Indicates key points From the perspective The current frame light and The distance; and Representative key points From different perspectives and perspective The light in the middle; Indicates key points From the perspective Perspective The geometric consistency metric between them, where T represents the transpose operation;

[0040] 3D world coordinates and 2D pose The objective function for optimizing camera extrinsic parameters is constructed as follows:

[0041] ;

[0042] ;

[0043] In the formula, Camera extrinsic parameters representing N viewpoints, From the perspective Camera extrinsic parameters; first data item For positional deviation, Data items to prevent artifacts from multi-user interaction; It is the Görmann-McLuhr function; The number of joints in a digital 3D human body; Indicates key points World coordinates obtained from a video frame at a certain perspective. Indicates the use of perspective Camera external parameters Will Coordinates projected onto the 2D image plane;

[0044] The two constructed constraints are combined to form physical geometric consistency. Image frames that do not meet the physical geometric consistency constraints are filtered out, and the objective function for optimizing the camera extrinsic parameters is solved using the image frames that meet the physical geometric consistency constraints, resulting in camera extrinsic parameters for N viewpoints. .

[0045] Furthermore, data items to prevent multi-user interaction artifacts It is calculated based on the differentiable symbolic distance field and is expressed as:

[0046] ;

[0047] In the formula, For human body sampling points, For the first A digital 3D human body mesh surface. This serves as an index for different human figures within a video frame. The number of human figures in the video frame; Sampling points To mesh surface The distance.

[0048] Furthermore, the process of obtaining the final digital 3D human motion through multi-view fitting specifically includes:

[0049] (1) After transforming all perspectives to the world coordinate system, the parameters of the digital three-dimensional human body are averaged to obtain a global digital three-dimensional human body;

[0050] (2) Using the parameters of each camera, the currently obtained global digital 3D human body is reprojected onto each view;

[0051] (3) Use the human body 3D reconstruction model to perform human body 3D reconstruction on each view to obtain digital 3D human body in each camera coordinate system;

[0052] (4) Using the extrinsic parameters of each camera, the digital 3D human body in the corresponding camera coordinate system is transformed to the world coordinate system, and then the process is repeated until the preset number of iterations is reached.

[0053] A multi-view, uncalibrated human motion capture system includes a memory and a processor. The memory stores a computer program, and when the computer program is executed by the processor, the processor enables the processor to implement any of the aforementioned multi-view, uncalibrated human motion capture methods.

[0054] This invention relates to a multi-view, uncalibrated method and system for rapid human motion capture. First, it utilizes a trained single-view 3D human reconstruction model and an existing 2D pose tracking framework to obtain the motion trajectories of the human skeleton's joints from multiple perspectives. Then, it uses physical geometric consistency to constrain the digital 3D human figure from each perspective, continuously optimizing the camera's extrinsic parameters to achieve calibration. Finally, it employs multi-view fitting to output the digital 3D human motion. Compared with existing visual human motion capture technologies, this invention has the following advantages:

[0055] (1) By directly capturing key points of the human body and digital three-dimensional human body through computer vision algorithms, no heavy calibration work is required. Just fix the camera position and record a video frame that can capture human body activities from various angles to achieve automatic calibration.

[0056] (2) In the single-view human three-dimensional reconstruction model, the YOLO-adapted feature extraction network is used as the backbone network of the model and multiple gated neural units (GRU) are used as updaters, which combines real-time performance and high accuracy.

[0057] (3) The single-view human three-dimensional reconstruction model is applied to multiple views through multi-view fitting to realize multi-view human motion capture and human reconstruction. When a certain view is wrong, the present invention can still output digital three-dimensional human motion with higher accuracy than single view. Attached Figure Description

[0058] Figure 1 This is a flowchart illustrating the rapid capture of human motion from multiple angles without calibration, according to an embodiment of the present invention.

[0059] Figure 2 This is a schematic diagram of a single-view human body 3D reconstruction model module according to an embodiment of the present invention.

[0060] Figure 3 This is a schematic diagram of the camera calibration process according to an embodiment of the present invention.

[0061] Figure 4 This is a reasoning diagram of a multi-view human body 3D reconstruction model according to an embodiment of the present invention.

[0062] Figure 5 The image shows the results of multi-view digital three-dimensional human body reconstruction according to an embodiment of the present invention. (a1) to (a3) ​​are images from three different perspectives in the first time period, and (a4) is the digital three-dimensional human body reconstructed in the first time period; (b1) to (b3) are images from three different perspectives in the second time period, and (b4) is the digital three-dimensional human body reconstructed in the second time period; (c1) to (c3) are images from three different perspectives in the third time period, and (c4) is the digital three-dimensional human body reconstructed in the third time period. Detailed Implementation

[0063] The technical solution of this invention is based on a three-dimensional human body reconstruction model to calibrate the camera and capture human movements. It can be tested using the Python programming language or applied in engineering using the C / C++ programming language.

[0064] This embodiment provides a method for rapid human motion capture based on multi-view, uncalibrated motion capture. The processing flow is described in [link to documentation]. Figure 1 This includes the following steps.

[0065] Step 0: Build a neural network-based 3D human reconstruction model and train it on multiple publicly available single-view human datasets to ensure that the inference speed is improved as much as possible while achieving high accuracy.

[0066] The digital 3D human body in this embodiment adopts a skinned multi-person linear model, whose parameters include the shape vector, posture vector and offset vector of several skeletal joints of the human body.

[0067] The human body 3D reconstruction model obtained by training based on neural networks in this embodiment includes: a feature extraction network, a multilayer perceptron, a reprojection module, and an update module.

[0068] (1) Feature extraction network.

[0069] The feature extraction network, adapted from YOLO, serves as the backbone network of the human 3D reconstruction model, extracting features from images captured from multiple viewpoints. The network employs a multi-scale feature fusion mechanism and an attention enhancement mechanism. Its overall architecture includes efficient multi-feature extraction layers and hierarchical feature fusion paths. Through reparameterization, it balances model accuracy and inference speed. Figure 2 As shown.

[0070] The multi-feature extraction layer adopts a three-stage hierarchical structure, constructing a multi-scale feature space through progressive downsampling: Shallow feature extraction consists of 5 convolutional modules, including two 3×3 standard convolutions and three C3k2 modules, outputting a 512-dimensional feature map. This stage preserves rich spatial detail information, suitable for initial human contour localization; Mid-level feature extraction obtains 1024-dimensional features through two downsampling passes, employing C3k2 modules with residual connections to enhance feature reuse efficiency; Deep semantic extraction introduces an SPPF spatial pyramid pooling module to expand the receptive field, combined with a dual-channel C2PSA attention module to enhance keypoint response features. The final output is a 1024-dimensional feature map F containing high-dimensional semantic information.

[0071] In this embodiment, the C3k2 module is a dual-convolution fast residual feature extraction module, the SPPF module is a fast spatial pyramid pooling module, and the C2PSA attention module is an efficient feature extraction module with fused attention. These are all well-known professional names in neural networks.

[0072] (2) Multilayer perceptron.

[0073] Multilayer perceptron (MLP) unwraps the feature map F extracted by the feature extraction network to obtain a coarse-grained digital 3D human body.

[0074] (3) Reprojection module.

[0075] The reprojection module maps the joints of a coarse-grained digital 3D human body to the feature map F extracted by the feature extraction network for spatial feature retrieval. Each joint is mapped to a single feature channel, resulting in K channels corresponding to K joints. This design aims to allow the feature extraction layer to acquire feature information from different joints.

[0076] (4) Update module.

[0077] The update module is constructed using 49 gated neural units (GRUs) and a regression multilayer perceptron. Each joint uses one gated neural unit to update the coarse-grained digital 3D human body and obtain a fine-grained digital 3D human body, thus achieving single-view 3D reconstruction of the human body.

[0078] For key points Reprojection points obtained by reprojecting onto the corresponding channel feature map (u, v are pixel coordinates), using Centered window features:

[0079] (1);

[0080] In the formula: r = 3 pixels is the window radius. This is achieved by integrating feedback information from all key points and combining it with the estimated parameters of the current frame t. Boundary box center and scale Construct the final feedback vector:

[0081] (2);

[0082] The feature map for each joint does not directly detect a single joint, but instead learns features associated with that joint. Based on the number of joints in the skinned multi-person linear model (SMPL), 49 gated recurrent units (GRUs) and one regressive multilayer perceptron were used as the update module. The update module feeds back the signal... (As shown in Equation 2) is used as input to predict the correction amount of the digital 3D human body. Add it To generate a digital 3D human body with refined coarse granularity:

[0083] (3);

[0084] Each gated neural unit is responsible for updating the SMPL parameters of a joint, thereby refining the coarse-grained digital 3D human body.

[0085] In this embodiment, the loss function for training the human body 3D reconstruction model includes three parts: 3D joint loss, 2D projection joint loss, and SMPL model update loss.

[0086] (1) Loss of 3D joints. This loss function is used to calculate the difference between the predicted joints and the ground truth joints. The L1 loss function is used. The calculation formula is:

[0087] (4);

[0088] In the formula, The true value of a 3D joint is determined by wearing sensors to determine its position in space; These are 3D joints of a digital 3D human body, representing coarse-grained and fine-grained methods, respectively. It refers to the number of joints.

[0089] (2) Two-dimensional projection joint loss: To achieve alignment between the reconstructed digital 3D human body and the input 2D image, projection alignment is achieved by estimating the parameters of the weak perspective camera. The calculation formula is as follows:

[0090] (5);

[0091] In the formula, These are the 2D joints of a digital 3D human body with coarse and fine granularity, respectively, as viewed from the current perspective.

[0092] (3) SMPL model update loss, which is calculated from the coarse-grained SMPL model The fine-grained SMPL model is obtained after the update module. The formula for calculating the difference in change is as follows:

[0093] (6);

[0094] The following hyperparameters were used during training: the optimizer employed SGD with a momentum of 0.937 and a weight decay factor of 0.0005. The batch size was 64, the initial learning rate (LRO) was 0.001, and the training consisted of 100 iterations. A multi-stage learning rate was used, with the learning rate updated at a decay factor of 0.1. Requires Ubuntu 22.04LTS or later, with PyTorch 1.11 and Python 3.9 or later. The hardware platform included two NVIDIA GeForce RTX 3090 graphics cards, CUDA 11.8, and CuDNN8700, along with a CPU, at least 64GB of RAM, and a solid-state drive of at least 1TB.

[0095] Step 1: Using a human body 3D reconstruction model trained based on a neural network, perform corresponding human body 3D reconstruction from different perspectives based on videos from different angles, and obtain digital 3D human body motion in each camera coordinate system.

[0096] This embodiment uses SMPL to represent human motion, which consists of a series of parameters: shape vector Attitude vector and offset vector .

[0097] The human body 3D reconstruction model of this invention takes image frames as input and outputs corresponding digital 3D human bodies. A video composed of image frames from multiple consecutive time points is used to output digital 3D human body motion from the corresponding output of the human body 3D reconstruction model.

[0098] We use an existing 2D pose estimation and tracking framework to obtain the tracking 2D pose of each person in this video frame, and use the human 3D reconstruction model from step 1 to obtain SMPL models with different viewpoints and timestamps.

[0099] Step 2: Estimate an initial extrinsic parameter for each camera from a different perspective. Then, using the initial extrinsic parameter of each camera and the digital 3D human motion in the corresponding camera coordinate system, and based on the physical geometry consistency strategy, iteratively calculate the extrinsic parameter of each camera.

[0100] Step 2.1, estimate an initial extrinsic parameter for each camera from a different viewpoint, including:

[0101] First, using video frames from various viewpoints, existing 2D pose estimation and tracking methods are employed to estimate the 2D pose of the human body and its corresponding confidence level in each viewpoint; then, the joints are... From a perspective The 2D pose and confidence obtained from the video frames are denoted as follows: , .

[0102] Then, using the epipolar geometry of the known intrinsic parameters of the camera, the extrinsic parameters of each camera are solved from the 2D pose of the human body from multiple perspectives, and used as the initial extrinsic parameters.

[0103] Step 2.2: Using the initial extrinsic parameters of each camera and the digital 3D human motion in the corresponding camera coordinate system, and based on the physical geometry consistency strategy, iteratively calculate the extrinsic parameters of each camera.

[0104] Step 2.2.1, construct the physical-geometric consistency constraints, refer to... Figure 3 As shown.

[0105] First, using the initial extrinsic parameters of each camera, the digital 3D human body in the corresponding camera coordinate system is transformed to the world coordinate system; then the joints are... From a perspective The world coordinates obtained from the video frames are as follows ;

[0106] Then, considering that although the camera parameters are inaccurate, the triangulated skeleton joint trajectories obtained from the 2D pose with precise identity are theoretically continuous, a set of rays from the camera optical center and passing through the corresponding 3D joint coordinates is used to construct physical constraints. Based on the distance from the joint point's position in the previous frame to the ray in the current frame at the same viewpoint, the physical consistency constraints are constructed as follows:

[0107] ;

[0108] Then, the rays from the joints in different viewpoints are forced to share a common ground, and a geometric consistency constraint is constructed as follows:

[0109] ;

[0110] In the formula, Representative key points From the perspective The light in the middle, It is a direction vector. This is the moment vector of the light ray, which is used to encode the position information of the light ray in space; Key points From the perspective The world coordinates relative to the previous frame t-1 in the current frame; Indicates key points From the perspective The current frame light and The distance; and Representative key points From different perspectives and perspective The light in the middle; Indicates key points exist Perspective The geometric consistency metric between them, where T represents the transpose operation.

[0111] In this embodiment, for the view For each detected 2D joint coordinate, a 3D ray is generated by reverse-engineering the camera projection model. This ray originates from the camera's optical center, passes through the 2D joint pixel position, and extends into 3D space. The mathematical representation of this ray... , , used to describe the positional relationship of light rays in space.

[0112] This embodiment uses the multi-view of the first frame. Figure 2 The initial 3D joint positions are triangulated by combining the pose detection results with the initial camera parameters and minimizing the reprojection error. Subsequent frames (t≥1) are obtained through iterative calculation.

[0113] Step 2.2.2, based on 3D world coordinates and 2D pose The objective function for optimizing camera extrinsic parameters is constructed as follows:

[0114] ;

[0115] ;

[0116] ;

[0117] In the formula, Camera extrinsic parameters representing N viewpoints, It's a perspective The corresponding camera extrinsic parameters include rotation and translation; the first data item. For positional deviation, Data items to prevent artifacts from multi-user interaction; It is the Görmann-McLuhr function; The number of joints in a digital 3D human body; For human body sampling points, For the first A digital 3D human body mesh surface. This serves as an index for different human figures within a video frame. The number of human figures in the video frame; Sampling points To mesh surface The distance.

[0118] Step 2.2.3 combines the two constraints established in Step 2.2.1 to form physical geometric consistency, filtering out image frames that do not satisfy the physical geometric consistency constraints, thus subjecting the entire human body movement to inherent kinematic and dynamic constraints; then, using the image frames with physical geometric consistency constraints, the objective function for optimizing the camera extrinsic parameters is solved, resulting in camera extrinsic parameters for N viewpoints. .

[0119] Due to the inherent ambiguity of human-to-human interactions and occlusion, existing pose detection and tracking methods struggle to obtain accurate 2D poses with precise identities from live video. Drift and jitter generated by pose detection are often high-frequency, while identity errors generated by pose tracking are low-frequency. The mixing of these two types of noise is well-known in multi-person mesh reconstruction. To address this obstacle, this embodiment constructs physical-geometric consistency constraints to simultaneously reduce high-frequency and low-frequency noise in 2D pose from each view, removing erroneous and inaccurate image frames and enabling more precise camera extrinsic parameter optimization.

[0120] Furthermore, the camera parameters obtained from the above iterative calculations were qualitatively and quantitatively evaluated. Due to the rigid transformation between the predicted camera parameters and the ground truth provided by the dataset, rigid registration was performed on the estimated camera. As shown in Table 1, the results were evaluated using position error, angle error, and reprojection error. Under relatively large-scale views, the method of this invention exhibits good accuracy across all metrics.

[0121] ;

[0122] Step 3: Using the camera extrinsic parameters obtained through iterative calculation, the digital 3D human body in the corresponding camera coordinate system is transformed to the world coordinate system, and then the final digital 3D human body motion is obtained through multi-view fitting.

[0123] Among them, the final digital 3D human motion is obtained through multi-view fitting, with reference to Figure 4 As shown, it includes:

[0124] (1) The shape vector β and posture vector of the digital 3D human body after all perspectives are transformed to the world coordinate system The average values ​​are taken respectively to obtain a global digital 3D human body.

[0125] In multi-view merging, the handling of β relies heavily on PCA properties to achieve efficient merging. This is achieved through a simple averaging operation: in each iteration, the model independently processes the image for each viewpoint, predicting a set of shape parameters β for each viewpoint. These β values ​​are then merged through a "multi-view averaging process," which takes the average of the PCA coefficients predicted from all views. , It's the number of viewpoints. It is the first The shape vectors are obtained from the prediction of each viewpoint.

[0126] Attitude vector To represent joint rotation, a "rotation average" is used during merging. The specific steps are: calculate the average of the rotation matrix for all viewpoints; then use SVD (singular value decomposition) to project the average result back to the SO(3) space (special orthogonal group) to ensure the effectiveness and smoothness of the rotation.

[0127] In the SMPL model, the translation parameters of the root node represent the overall position of the model in 3D space. Its multi-view merging process involves using camera extrinsic parameters to combine the estimated translation parameters from each viewpoint. By unifying the estimates to the world coordinate system and performing a (weighted) average of the estimated values ​​in the unified coordinate system, a global estimate is obtained. The weighted average is a weighted average based on the estimated confidence scores of each viewpoint. The confidence scores are derived from the variance of the model's predictions, the proximity of the viewpoint to the subject (e.g., less occlusion, higher confidence scores for frontal views), etc.

[0128] (2) Using the parameters of each camera, the currently obtained global digital 3D human body is reprojected onto each view.

[0129] (3) Use the human body 3D reconstruction model to perform human body 3D reconstruction on each view to obtain the digital 3D human body in each camera coordinate system.

[0130] (4) Using the extrinsic parameters of each camera, the digital 3D human body in the corresponding camera coordinate system is transformed to the world coordinate system, and the process is repeated in step (1) until the preset number of iterations is reached.

[0131] like Figure 5 As shown, the three perspectives are reconstructed by a single-view human body 3D reconstruction model and then output as digital 3D human body motion through multi-view fitting.

[0132] To verify the effectiveness and efficiency of the model and method, a comparison was made on a self-built experimental dataset and several different single-view human 3D reconstruction models, as shown in Table 2. The model outperformed other methods in all performance metrics.

[0133] ;

[0134] The above embodiments are preferred embodiments of this application. Those skilled in the art can make various changes or improvements based on them. Without departing from the overall concept of this application, these changes or improvements should fall within the scope of protection claimed in this application.

Claims

1. A method for rapid human motion capture based on multi-view, uncalibrated imaging, characterized in that, include: Using a neural network-based and trained 3D human body reconstruction model, 3D human body reconstruction is performed from the perspective of the video from different angles to obtain digital 3D human body motion in each camera coordinate system. An initial extrinsic parameter is estimated for each camera from a different perspective. Then, using the initial extrinsic parameter of each camera and the digital 3D human motion in the corresponding camera coordinate system, and based on a physical-geometric consistency strategy, the extrinsic parameter of each camera is iteratively calculated. Using the camera extrinsic parameters obtained through iterative calculation, the digital 3D human motion in the corresponding camera coordinate system is transformed to the world coordinate system, and then the final digital 3D human motion is obtained through multi-view fitting. The estimation of an initial extrinsic parameter for each camera from a different viewpoint includes: First, using video frames from various viewpoints, existing 2D pose estimation and tracking methods are employed to estimate the 2D pose of the human body and its corresponding confidence level in each viewpoint; then, the joints are... From a perspective The 2D pose and confidence obtained from the video frames are denoted as follows: , ; Then, using the epipolar geometry of the known intrinsic parameters of the camera, the extrinsic parameters of each camera are solved from the multi-view human 2D pose and used as the initial extrinsic parameters. The iterative calculation of the extrinsic parameters of each camera includes: Using the initial extrinsic parameters of each camera, the digital 3D human body in the corresponding camera coordinate system is transformed to the world coordinate system; the joints are... From a perspective The world coordinates obtained from the video frames are as follows ; Based on the distance between the position of the key point in the previous frame and the light ray in the current frame from the same viewpoint, the physical consistency constraint is constructed as follows: ; By forcing the rays from the joints to coexist at different viewpoints, a geometric consistency constraint is constructed as follows: ; In the formula, Representative key points From the perspective The light in the middle, It is a direction vector. This is the moment vector of the light ray, which is used to encode the position information of the light ray in space; Key points From the perspective The world coordinates relative to the previous frame t-1 in the current frame; Indicates key points From the perspective The current frame light and The distance; and Representative key points From different perspectives and perspective The light in the middle; Indicates key points From the perspective Perspective The geometric consistency metric between them, where T represents the transpose operation; 3D world coordinates and 2D pose The objective function for optimizing camera extrinsic parameters is constructed as follows: ; ; In the formula, Camera extrinsic parameters representing N viewpoints, From the perspective Camera extrinsic parameters; first data item For positional deviation, Data items to prevent artifacts from multi-user interaction; It is the Görmann-McLuhr function; The number of joints in a digital 3D human body; Indicates key points World coordinates obtained from a video frame at a certain perspective. Indicates the use of perspective Camera external parameters Will Coordinates projected onto the 2D image plane; The two constructed constraints are combined to form physical geometric consistency. Image frames that do not meet the physical geometric consistency constraints are filtered out, and the objective function for optimizing the camera extrinsic parameters is solved using the image frames that meet the physical geometric consistency constraints, resulting in camera extrinsic parameters for N viewpoints. .

2. The method for rapid human motion capture according to claim 1, characterized in that, The digital 3D human body adopts a skinned multi-person linear model, whose parameters include the shape vector, posture vector and offset vector of several skeletal joints of the human body.

3. The method for rapid human motion capture according to claim 2, characterized in that, The human 3D reconstruction model based on and trained a neural network includes: The feature extraction network employs YOLO's multi-scale feature fusion mechanism and attention enhancement to extract features from the input video frames; Multilayer perceptron is used to reconstruct coarse-grained digital 3D human bodies based on extracted features. The reprojection module is used to map the digital 3D human body onto the extracted features to obtain a feedback vector; The update module, including a gated neural network and a regressive multilayer perceptron, is used to unwrap and regress the feedback vector to obtain a fine-grained digital 3D human body.

4. The method for rapid human motion capture according to claim 3, characterized in that, The feature extraction network includes multiple feature extraction layers and hierarchical feature fusion paths; The multi-feature extraction layer adopts a three-stage hierarchical structure and constructs a multi-scale feature space through progressive downsampling: (1) shallow feature extraction, including two 3×3 standard convolutions and three C3k2 modules, outputting a 512-dimensional feature map; (2) mid-level feature extraction, obtaining 1024-dimensional features through two downsamplings, and using a C3k2 module with residual connections to enhance feature reuse efficiency; (3) deep semantic extraction, introducing an SPPF spatial pyramid pooling module to expand the receptive field, and cooperating with a dual-channel C2PSA attention module to enhance key point response features; finally, a 1024-dimensional feature map containing high-dimensional semantic information is output.

5. The method for rapid human motion capture according to claim 3, characterized in that, The loss functions used to train the human 3D reconstruction model include: 3D joint loss, 2D projection joint loss, and SMPL model update loss; among which; The three-dimensional joint loss is denoted as The calculation formula is: ; in, For the true values ​​of the three-dimensional joints, These are 3D joints of a digital 3D human body, representing both coarse-grained and fine-grained types. It is the number of joints; The two-dimensional projection joint loss is denoted as The calculation formula is: ; In the formula, This represents the mapping of the true values ​​of 3D joints to the current viewpoint. These are 2D joints of a digital 3D human body with coarse and fine granularity, respectively, viewed from the current perspective. The SMPL model update loss is denoted as... The calculation formula is as follows: ; In the formula, and These are digital 3D human figures with coarse and fine granularity, respectively.

6. The method for rapid human motion capture according to claim 1, characterized in that, Data items to prevent multi-user interaction artifacts It is calculated based on the differentiable symbolic distance field and is expressed as: ; In the formula, For human body sampling points, For the first A digital 3D human body mesh surface. This serves as an index for different human figures within a video frame. The number of human figures in the video frame; Sampling points To mesh surface The distance.

7. The method for rapid human motion capture according to claim 1, characterized in that, The process of obtaining the final digital 3D human motion through multi-view fitting specifically includes: (1) After transforming all perspectives to the world coordinate system, the parameters of the digital three-dimensional human body are averaged to obtain a global digital three-dimensional human body; (2) Using the parameters of each camera, the currently obtained global digital 3D human body is reprojected onto each view; (3) Use the human body 3D reconstruction model to perform human body 3D reconstruction on each view to obtain digital 3D human body in each camera coordinate system; (4) Using the extrinsic parameters of each camera, the digital 3D human body in the corresponding camera coordinate system is transformed to the world coordinate system, and then the process is repeated until the preset number of iterations is reached.

8. A multi-view, uncalibrated human motion capture system, comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the computer program is executed by the processor, it causes the processor to implement the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Pipeline three-dimensional reconstruction and pit quantification method based on multi-view geometry

    CN116363302A

  • Monocular video-based multi-stage human motion capture method and device, and medium

    CN116386141A