Markerless monocular video 3d pose reconstruction method and system

CN122550765APending Publication Date: 2026-08-11COMMUNICATION UNIVERSITY OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-15
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

但在高动态动作场景中,模糊帧可能占据连续数帧甚至十余帧,剔除会导致最具价值的动作段落完全缺失

Benefits of technology

[0025]第一,改善了长期以来三维重建和姿态估计领域将运动模糊帧作为无效数据予以丢弃的技术共识,将模糊帧转化为高时间分辨率运动信息的有效载体。在高动态动作区间等效将24帧/秒提升至96至240帧/秒的姿态采样率,使原本完全不可用的帧转变为比清晰帧更密集的运动采样。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122550765A_ABST
    Figure CN122550765A_ABST
Patent Text Reader

Abstract

This invention discloses a label-free monocular video 3D pose reconstruction method and system for motion capture, relating to the fields of computer vision and 3D pose estimation. The method constructs a motion field by performing regional fuzzy kernel estimation on blurred frames, expands the blurred frames into multiple sub-poses to increase the sampling rate, and applies a triple joint iterative optimization by applying optical forward consistency constraints, inverse dynamic ground reaction force physical constraints, and motion field-dynamic cross-validation constraints. During iteration, contact states are dynamically determined, and physical constraint feedback corrects the fuzzy kernel parameters to form a closed loop. This invention is applicable to low-cost label-free motion capture in high-dynamic motion scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision and 3D pose estimation technology, specifically to a method and system for 3D pose reconstruction from motion capture-free monocular video. Background Technology

[0002] In film and television special effects and animation production, actor motion capture traditionally relies on optical marker motion capture studios or inertial motion capture suits. Optical marker methods require attaching dozens to hundreds of reflective markers to the actor's body, with multiple infrared cameras performing triangulation within a dedicated motion capture studio. This method is space-consuming, has complex equipment setup, and the dense markers restrict the actor's movement, affecting the naturalness of the performance. Inertial motion capture suits collect angular velocity and acceleration information through inertial measurement units distributed throughout the body, but suffer from accumulated inertial drift errors and magnetic field interference, resulting in significant degradation in positioning accuracy in live-action shooting environments. Both traditional methods require substantial equipment investment and dedicated space configurations, making them unaffordable for independent film, short drama, advertising, and game cutscene production teams due to budget constraints.

[0003] Deep learning-based monocular video 3D pose estimation technology can directly infer the 3D spatial coordinates of key points of the human skeleton from videos shot by a single ordinary camera, without the need for marker points or special locations. However, existing monocular 3D pose estimation methods face three overlapping core technical challenges in high-dynamic action scenes in film and television shooting.

[0004] The first challenge is the loss of visual information caused by motion blur. Standard cinema cameras typically operate at 24 to 25 frames per second, while in high-dynamic-range action, the linear velocity of limb extremities can reach 8 to 12 meters per second. At a standard shutter speed, the end-joints create a 30 to 60 pixel blur in a single frame. Keypoint detection networks completely collapse under this level of blur, rendering the 3D pose of the blurred frames completely unusable. Existing methods for handling motion-blurred frames either discard or tolerate them—the long-standing consensus in 3D reconstruction and pose estimation considers motion-blurred frames as invalid data that impairs accuracy and should be detected and discarded. However, in high-dynamic-range action scenes, blurred frames may occupy several or even more than ten consecutive frames; discarding them would result in the complete loss of the most valuable action sequences.

[0005] The second challenge is the insufficient numerical basis of physical constraints. High-dynamic movements involve frequent transitions between airborne and contact states. Purely data-driven regression models lack prior knowledge of Newtonian mechanics, and the output 3D posture may exhibit physical inconsistencies such as feet not leaving the ground during the airborne phase, the body penetrating the ground during the landing phase, and the supporting foot sliding horizontally during the walking phase. Although there have been attempts to apply inverse dynamics physical constraints to monocular video posture optimization, these methods apply to posture sequences at the original frame rate (24 frames / second). The effective application of physical constraints depends on the second-order time derivative (acceleration) of the posture sequence. However, at a low sampling rate of 24 frames / second, the inter-frame displacement of high-speed moving joints reaches 15 to 20 centimeters. The numerical truncation error of second-order finite difference for such a large step size is extremely large—the relative error of acceleration can reach 30% to 50%—making the inverse dynamics reaction force estimation completely unreliable numerically, rendering the physical constraints ineffective.

[0006] The third challenge is the non-uniqueness of the optical constraint solution space. Even if a blurred frame is understood as the integral of light intensity across multiple instantaneous poses during exposure, and an optical forward consistency constraint is established through differentiable rendering that "sub-pose superposition should reproduce the actual blur," this constraint equation set remains highly underdetermined—multiple different sub-pose sequences can produce almost identical composite blurred images. Without additional constraints independent of the optical domain, optical consistency alone cannot yield a unique and physically plausible solution.

[0007] The three difficulties mentioned above are intertwined, forming a vicious cycle: the frames with the most severe motion blur correspond to the moments of most intense action, and are also the frames where the physical state transitions require the most accurate constraints; however, these frames are unusable due to blurring, and at the same time, physical constraints are also unusable due to insufficient numerical precision at the original frame rate; while optical constraints can be established, the solution is not unique. Summary of the Invention

[0008] Technical objective: To address the triple failure dilemma of "lack of visual information, insufficient numerical basis of physical constraints, and non-unique solutions of optical constraints" in existing high-dynamic motion scenarios, this invention discloses a method and system for reconstructing three-dimensional pose of motion capture unmarked monocular video.

[0009] The core technical concept of this invention is to treat motion-blurred frames not as invalid data to be discarded, but as integral information carriers of multiple instantaneous postures during exposure. By extracting the motion direction and velocity information of each segment through regional blur kernel analysis, a motion field is constructed. The blurred frame is expanded into multiple sub-postures, effectively increasing the temporal sampling rate. Simultaneously, a triple joint optimization framework is established, consisting of optical forward consistency constraints, inverse dynamics virtual ground reaction force physical constraints, and motion field-dynamics cross-validation constraints. These three constraints shrink the solution space to a near-unique solution from three independent directions: the two-dimensional imaging domain, the three-dimensional mechanical domain, and the kinematic velocity domain. The increased sampling rate reduces the relative error of the numerical differential of acceleration from 30% to 50% at the original frame rate to 3% to 5%, enabling the inverse dynamics physical constraints to have a usable numerical basis for the first time in high-dynamic scenes. Furthermore, a closed loop is established to correct the blur kernel parameters based on physical constraints—when the velocity derived from inverse dynamics is inconsistent with the velocity derived from the initial blur kernel, the blur kernel parameters are corrected, forming a two-way closed loop between image analysis and physical verification.

[0010] Technical solution: To achieve the above technical objectives, the present invention adopts the following technical solution:

[0011] This invention provides a method for reconstructing 3D pose from motion-captured, markerless monocular video, comprising the following steps:

[0012] The motion blur level of each frame of the monocular video is evaluated at the frame level, and each frame is divided into clear frames and blurred frames based on the evaluation results.

[0013] Perform 3D pose estimation on clear frames to obtain the 3D pose and global human body morphology parameters corresponding to the clear frames;

[0014] The blurred frame is subjected to regional blur kernel estimation. The human body region is divided into multiple sub-regions corresponding to body parts and the blur kernel parameters of each sub-region are estimated independently to construct the whole body motion field. The number of sub-frames is dynamically determined according to the blur kernel parameters.

[0015] Using the 3D pose of the clear frame as the temporal boundary condition, and combining the whole-body motion field, the blurred frame is initialized with sub-pose to generate multiple sub-poses.

[0016] A triple joint iterative optimization is performed by simultaneously applying optical forward consistency constraints, dynamic physical rationality constraints, and motion field-dynamic cross-validation constraints to multiple sub-poses. The optical forward consistency constraint is an image consistency constraint between the simulated blurred image synthesized from multiple sub-poses and the actual blurred frame image. The dynamic physical rationality constraint is a physical rationality constraint applied to the virtual ground reaction force obtained through inverse dynamics analysis based on global human morphology parameters and the motion sequence of multiple sub-poses. The motion field-dynamic cross-validation constraint is a consistency constraint between the inter-frame velocities of multiple sub-poses and the motion velocities of corresponding body parts in the whole-body motion field. Physically consistent sub-pose sequences are obtained through this triple joint iterative optimization.

[0017] The 3D pose of the clear frame and the sub-pose sequence obtained by unfolding the blurred frame are temporally stitched together, and the 3D pose reconstruction result is output after global dynamic consistency optimization.

[0018] This invention also provides a motion capture label-free monocular video 3D pose reconstruction system corresponding to the above-mentioned motion capture label-free monocular video 3D pose reconstruction method, including a frame quality assessment module, a basic pose estimation module, a fuzzy kernel motion field analysis module, a physical constraint sub-pose unfolding engine, and a global temporal optimization module.

[0019] The frame quality assessment module is used to assess the motion blur level of each frame of a monocular video and classify each frame into clear frames and blurred frames based on the assessment results.

[0020] The basic pose estimation module is used to perform 3D pose estimation on clear frames and obtain the 3D pose and global human body shape parameters corresponding to the clear frames.

[0021] The fuzzy kernel motion field analysis module is used to perform regional fuzzy kernel estimation on fuzzy frames. It divides the human body region into multiple sub-regions corresponding to body parts and independently estimates the fuzzy kernel parameters for each sub-region to construct the whole-body motion field. It also dynamically determines the number of sub-frames to be unfolded based on the fuzzy kernel parameters.

[0022] The physical constraint sub-pose unfolding engine uses the 3D pose of a clear frame as a temporal boundary condition, combined with the whole-body motion field to initialize and generate multiple sub-poses. It simultaneously applies optical forward consistency constraints, dynamic physical rationality constraints, and motion field-dynamic cross-validation constraints to multiple sub-poses for triple joint iterative optimization to obtain a physically consistent sub-pose sequence. The optical forward consistency constraint is the image consistency constraint between the simulated blurred image synthesized from multiple sub-poses and the actual blurred frame image. The dynamic physical rationality constraint is the physical rationality constraint applied to the virtual ground reaction force obtained by inverse dynamics analysis based on global human body morphology parameters and the motion sequence of multiple sub-poses. The motion field-dynamic cross-validation constraint is the consistency constraint between the inter-frame velocity of multiple sub-poses and the motion velocity of the corresponding body parts in the whole-body motion field.

[0023] The global temporal optimization module is used to temporally stitch together the 3D pose of the clear frame and the sub-pose sequence obtained by unfolding the blurred frame, and output the 3D pose reconstruction result after global dynamic consistency optimization.

[0024] Beneficial Effects: The method and system for markerless monocular video 3D pose reconstruction based on motion capture provided by this invention have the following beneficial effects:

[0025] First, it improves upon the long-standing technical consensus in the field of 3D reconstruction and pose estimation to discard motion-blurred frames as invalid data, transforming them into effective carriers of high temporal resolution motion information. In high-dynamic motion ranges, it effectively increases the pose sampling rate from 24 frames per second to 96 to 240 frames per second, turning previously unusable frames into more densely sampled motion data than clear frames.

[0026] Second, the increased sampling rate reduced the relative error of the numerical derivative of acceleration from 30% to 50% to 3% to 5%, enabling ground reaction force estimation based on inverse dynamics to achieve usable accuracy for the first time in high-dynamic motion scenarios. This causal chain of "utilization of motion blur → increased sampling rate → improved accuracy of acceleration derivative → physical constraints changing from infeasible to feasible" reveals a previously undiscovered intrinsic connection between the informational value of motion-blurred frames and the numerical basis of physical constraints.

[0027] Third, the optical forward consistency constraint, the dynamic physical rationality constraint, and the motion field-dynamic cross-validation constraint constrain the sub-attitude from three independent directions: the two-dimensional imaging domain, the three-dimensional mechanical domain, and the kinematic velocity domain, respectively. The constraint directions of the three are nearly orthogonal, and their intersection shrinks the highly underdetermined expanded solution space to a near-unique solution. Compared with the scheme using only optical constraints, the triple constraint eliminates the non-uniqueness of the solution; compared with the scheme using only optical and physical constraints, the motion field-dynamic cross-validation constraint provides an additional velocity domain verification dimension.

[0028] Fourth, the closed loop of physical constraint feedback to correct fuzzy kernel parameters forms a two-way information flow of "image analysis → physical verification → image correction"—local errors in the initial fuzzy kernel estimation (such as the orientation angle deviation caused by insufficient signal-to-noise ratio in a certain sub-region) can be automatically corrected through physical constraint feedback without manual intervention, which significantly enhances the robustness of the system to the initial estimation accuracy of the fuzzy kernel.

[0029] Fifth, the system dynamically determines the take-off / contact state and the contact body part during iteration (expanding from the traditional two points on the soles of the feet to six types of parts: hands, shoulders, back, buttocks, knees, and soles of the feet), so that the precise moment of take-off and landing automatically emerges during the optimization process, with positioning accuracy reaching the sub-frame level, adapting to full-body rolling contact movements such as somersaults and rolls.

[0030] Sixth, the system only requires a single ordinary camera as an input device, without the need for marker points, dedicated motion capture sheds, or inertial sensors. Attached Figure Description

[0031] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.

[0032] Figure 1 This is a schematic diagram of the overall system architecture of the present invention;

[0033] Figure 2 Flowchart for frame quality assessment and classification;

[0034] Figure 3 A schematic diagram of regional fuzzy kernel estimation and whole-body motion field construction;

[0035] Figure 4 A flowchart of the triple joint iterative optimization of the physical constraint sub-pose unfolding engine;

[0036] Figure 5 A schematic diagram of differentiable rendering and synthetic blur comparison under optical forward consistency constraints;

[0037] Figure 6 A schematic diagram illustrating the virtual estimation of inverse dynamic ground reaction force and four physical rationality constraints;

[0038] Figure 7 A schematic diagram of the fuzzy kernel closed loop for cross-validation constraints and physical feedback of motion field-dynamics cross-validation;

[0039] Figure 8 A schematic diagram illustrating the dynamic determination and state machine transition of the airborne-contact state;

[0040] Figure 9 Output a flowchart for global timing optimization and sequence concatenation;

[0041] Figure 10 This is a schematic diagram showing the implementation effect of the method of the present invention applied to a somersault scenario. Detailed Implementation

[0042] The present invention will now be described more clearly and completely by way of a preferred embodiment in conjunction with the accompanying drawings, but this does not limit the invention to the scope of the described embodiment.

[0043] like Figure 1 As shown, the system of this invention includes five core modules: a frame quality assessment module, a basic pose estimation module, a fuzzy kernel motion field analysis module, a physically constrained sub-pose unfolding engine, and a global temporal optimization module. Monocular video is used as system input, and after being processed by the above modules, a physically consistent 3D pose sequence is output.

[0044] The system's data flow path is as follows: raw video frames first enter the frame quality assessment module for blur assessment and classification; frames classified as clear are sent to the basic pose estimation module to obtain 3D pose and global human body morphology parameters; frames classified as blurry are sent to the fuzzy kernel motion field analysis module to extract motion field information and determine the number of subframes; the pose of clear frames, as boundary conditions, and the motion field information, as guiding conditions, are jointly input into the physical constraint sub-pose unfolding engine for triple joint iterative optimization; the physical constraint sub-pose unfolding engine contains a feedback correction channel from physical constraints to the fuzzy kernel motion field analysis module; all pose data are merged into the global temporal optimization module to complete sequence splicing and global correction before output.

[0045] like Figure 2 As shown, the frame quality assessment module performs a quantitative assessment of the motion blur level of each frame of the input video.

[0046] For each frame, a lightweight human detector is first used to obtain the human body bounding box. Within the bounding box, two blur metrics for the human body region are calculated. The first metric is the Laplacian variance metric: the response map of the Laplacian operator is calculated for the human body region image, and the variance value of the response map is used as a sharpness measure. The second metric is the edge width metric: the main edges of the human body region are extracted, and the width of the gray-level gradient at the edges is calculated.

[0047] The two metrics are weighted and fused to obtain a blur severity score M for each frame, with a value ranging from 0 to 1. In this embodiment, the fusion formula is as follows: ,in Let Laplace variance be... The average edge width is represented by norm(·), which indicates normalization to the range of 0 to 1. and Set the values ​​to 0.6 and 0.4 respectively. Frames with an M value lower than 0.6 are classified as clear frames, and frames with an M value not lower than 0.6 are classified as blurry frames. This threshold can be adjusted according to specific shooting conditions.

[0048] The basic pose estimation module performs 3D pose estimation on clear frames. A multi-task pose estimation network is adopted, which takes a single frame of human image as input and jointly predicts 2D keypoint heatmap, 3D relative pose, global position of root node, and human morphological parameters.

[0049] The multi-task pose estimation network employs a shared feature extraction backbone with a multi-head output structure. The shared feature extraction backbone is a convolutional neural network containing multiple convolutional layers and residual connection modules. Four parallel output heads are set on top of the shared backbone: a 2D keypoint output head generates 2D heatmaps of each joint through deconvolutional layers; a 3D offset output head regresses the 3D coordinate offsets of each joint relative to the root node (hip center) through fully connected layers; a global position output head regresses the 3D position of the root node in the camera coordinate system through fully connected layers; and a morphological parameter output head regresses the human morphological parameter vector through fully connected layers.

[0050] The network's training data includes a public dataset with 3D annotations and a natural scene dataset with 2D annotations. The training loss function is a weighted sum of the losses from the four output heads: mean squared error loss for the 2D heatmap, L1 loss (minimum absolute deviation loss) for the 3D offset, L1 loss for the global location, and L2 regularization loss (minimum squared deviation regularization loss) for the morphological parameters. Training uses the Adam (Adaptive Moment Estimation) optimizer with an initial learning rate of 1×10⁻⁶. -3 The value is reduced to 0.1 times its original value in the 30th and 60th rounds respectively, with a total of 80 training rounds and a batch size of 64.

[0051] Human morphological parameters are shared globally throughout the video (the body shape of the same actor remains unchanged), and the median of the estimation results from all clear frames is taken as the global morphological anchor value. Global morphological parameters provide the basis for estimating the mass and moment of inertia of each limb segment: based on the estimated height H and weight W, the mass of each limb segment is calculated using the standard anthropometric regression equation. and moment of inertia The mass of each limb segment is distributed according to standard proportions: head 8.1% of total weight, trunk 49.7%, upper arm 2.8% (unilateral), forearm 1.6% (unilateral), hand 0.6% (unilateral), thigh 10.0% (unilateral), lower leg 4.7% (unilateral), and foot 1.4% (unilateral). The moment of inertia of each limb segment is approximated using a uniform cylinder based on the segment length and mass.

[0052] like Figure 3As shown, the fuzzy kernel motion field analysis module extracts the motion direction and speed information of each body part during the exposure of the fuzzy frame.

[0053] The human body region in the blurred frame was divided into 14 sub-regions corresponding to body parts: head, torso, left upper arm, left forearm, left hand, right upper arm, right forearm, right hand, left thigh, left calf, left foot, right thigh, right calf, and right foot. The sub-region division was based on the estimated 2D keypoint locations in the clear frame and the prior knowledge of human body part segmentation.

[0054] For each sub-region, the local blur kernel parameters—direction angle θ and blur length L—are estimated independently. For larger sub-regions (such as the torso, thighs, etc., with an area greater than a preset area threshold, which is 800 square pixels in this embodiment), the frequency domain zero-point fringe analysis method is used: a two-dimensional Fourier transform is performed on the sub-region image, and linear motion blur generates periodic zero-point fringes along the blur direction in the frequency domain. The direction of the zero-point fringes is the blur direction θ, and the reciprocal of the distance between adjacent zero points is the blur length L. For smaller sub-regions (such as hands, feet, etc., with an area not greater than the area threshold), the gradient direction histogram method is used: the gradient direction of all pixels in the sub-region is calculated and a histogram is plotted. Motion blur causes the gradient direction to concentrate in the direction perpendicular to the blur direction. The blur direction θ is inferred from the peak position of the histogram, and the blur length L is estimated from the spatial expansion range of the gradient magnitude.

[0055] The fuzz kernel parameters of the 14 sub-regions Assembled into a whole-body motion field, the following key information is extracted: the motion direction field, i.e., the motion direction of each limb segment during exposure; and the motion velocity field, i.e., the blur length of each limb segment. Divide by exposure time The obtained average velocity The motion consistency marker between parts refers to the motion field of adjacent limbs on the same kinematic chain satisfying kinematic constraints. Parts that do not satisfy these constraints are marked as low-confidence areas, and these markers will be corrected first in subsequent physical constraint feedback corrections.

[0056] The number N of subframes that need to be expanded from the blurred frame is dynamically determined based on the maximum blur length of the frame: ,in The maximum fuzzy length across all sub-regions. This is the preset maximum allowable joint displacement between subframes. In this embodiment... Set to 5 pixels. Taking a quick spinning kick as an example... When N=10 is approximately 50 pixels, this effectively increases the local sampling rate from 24 frames per second to 240 frames per second; taking a moderate-speed running motion as an example, When the subframes are approximately 15 pixels, N=3, effectively increasing the frame rate to 72 frames per second. Dynamically determining the number of subframes ensures that blurred frames with varying degrees of dynamic range receive appropriate unfolded resolution, avoiding computational waste due to over-unfolding or residual blurring due to under-unfolding.

[0057] Regarding the relationship between subframe time intervals and acceleration numerical accuracy: At the original 24 frames / second, the time interval between adjacent frames... Milliseconds. Central difference method for calculating acceleration. The truncation error is ,in Let be the linear acceleration vector of the joint at the k-th subframe. Let be the 3D position of the joint at the (k+1)th subframe. Let be the 3D position of the joint at the k-th subframe. This represents the time interval between subframes. The typical acceleration of the foot during a high-speed spinning kick is 100 m / s². 2 , The truncation error in milliseconds is approximately ±35m / s 2 (Relative error of 35%) makes the inverse dynamic reaction force estimation unreliable on an order of magnitude. Sub-pose unfolding distributes N=10 subframes evenly within the exposure window, with time intervals between subframes... In milliseconds (for a 20-millisecond exposure), the truncation error is reduced to approximately ±0.3 m / s. 2 (Relative error 0.3%). However, since the sub-attitude itself comes from optimization rather than direct measurement, its positional accuracy is limited by the convergence accuracy of optical and physical constraints. Therefore, the actual effective relative error is about 3% to 5%. This level of accuracy is sufficient to reliably distinguish between the two states of "airborne (GRF=0)" and "contact (GRF>0)", meeting the minimum numerical requirements for the effective application of physical constraints.

[0058] like Figure 4 As shown, the physical constraint sub-attitude unfolding engine is the core module of this invention, which realizes the triple joint iterative optimization of optical forward consistency constraints, dynamic physical rationality constraints, and motion field-dynamic cross-validation constraints, and also includes a closed loop for physical constraint feedback to correct fuzzy kernel parameters.

[0059] pose of the nearest clear frame before and after the blurred frame and As boundary conditions, initial values ​​for N sub-poses are generated through non-uniform interpolation, incorporating motion field information. .

[0060] Non-uniform interpolation strategy: For high-speed limbs with large blur lengths (such as the lower leg and foot during a kick), a larger displacement increment is allocated between subframes, with the increment direction determined by the blur kernel direction angle θ; for low-speed limbs with small blur lengths (such as the torso core), a smaller displacement increment is allocated between subframes. Specifically, the displacement increment of the i-th body part sub-region in the k-th subframe is related to the blur length of that part. Proportional: ,in Let be the displacement increment of the i-th body part sub-region at the k-th subframe. The sum of the fuzzy lengths of all body parts (normalized denominator). Let i be the 3D pose of the i-th body part sub-region in the nearest clear frame after the blurred frame. Let i be the 3D pose of the i-th body part sub-region in the nearest clear frame before the blurred frame. Let f be the time allocation function, which is linear under the assumption of uniform motion. This non-uniform initialization ensures that the initial trajectory of the high-speed segments roughly reflects the motion field constraints, while the low-speed segments maintain stable initial values ​​close to the boundary posture.

[0061] The total loss function for iterative optimization is defined as:

[0062]

[0063] in For optical forward consistency loss, For the loss of dynamic physical rationality, For the cross-validation loss of the motion field and dynamics, For time-series smoothing regularization loss, The joint angle biomechanical limit constraint loss is represented by ρ, β, γ, and κ, which are weighting coefficients for each component. In this embodiment, ρ = 0.5, β = 0.3, γ = 0.1, and κ = 0.05.

[0064] like Figure 5 As shown, optical forward coherence loss The calculation process is as follows: The parameters of N sub-poses are input into a differentiable human model renderer. This renderer accepts the 3D joint coordinates of the sub-poses and global human morphology parameters as input. Through differentiable mesh deformation and perspective projection operations, a 2D human contour-skeleton image corresponding to each sub-pose is generated. A composite blurred image is obtained by superimposing N two-dimensional images with uniform weights. .

[0065]

[0066] in For the actual blurred frame image, Measuring differences in edge blur patterns The edge loss weight is set to 0.3.

[0067] like Figure 6 As shown, the loss of dynamic physical rationality The calculation requires inverse dynamics recursion first.

[0068] Acceleration calculation: Linear acceleration is calculated using the central difference formula for the three-dimensional position of each joint in the sub-attitude sequence. ,in This is the time interval between subframes. The key point here is... The factor pose unfolds and is much smaller than the original frame interval. —With N=10 and For example, milliseconds milliseconds, only Approximately 1 / 21 of (42 milliseconds). Since the truncation error is O(Δt) 2 The 21-fold reduction in Δt decreases the truncation error by approximately 440 times. It is precisely this increased sampling rate achieved through sub-attitude expansion that provides the numerical accuracy foundation for inverse dynamics acceleration calculations, which was previously unattainable at the original frame rate.

[0069] Inverse dynamics recursion: Determining the mass of each limb segment using global human morphological parameters output by the basic pose estimation module. and moment of inertia The Newton-Euler equations are recursively applied from the distal limbs (hands, feet) towards the root node (hip): Applying the equations to each limb segment... Calculating inertial forces and their applications Calculate the moment of inertia, where F is the resultant force vector of the limb segment, m is the mass of the limb segment, and a is the acceleration vector along the center of mass of the limb segment. Let J be the resultant torque vector acting on the limb segment, and J be the moment of inertia tensor of the limb segment. The angular acceleration vector of the limb. The segmental angular velocity vector; after accumulating the internal forces and internal moments segment by segment to the root node, the remaining unbalanced forces minus gravity are the virtual ground reaction force (GRF) vector.

[0070] It is a weighted combination of four sub-items:

[0071] Consistency constraints of projectiles during takeoff In a subframe determined to be in an airborne state, the motion of the root node should satisfy the free projectile equation. ,in The root node's horizontal acceleration vector. Let g be the acceleration scalar in the vertical direction of the root node, and g be the acceleration due to gravity.

[0072] Friction cone constraint during contact phase In a subframe determined to be in contact, the tangential component of the GRF does not exceed the product of the normal component and the friction coefficient μ. ,in This is the tangential component vector of the virtual ground reaction force. For the normal component of the virtual ground reaction force scalar, Let μ be the coefficient of friction of the ground, and μ be 0.6.

[0073] Non-negative contact force constraint The normal component of GRF is non-negative. .

[0074] Rolling contact constraint of multiple parts of the body Dynamically detect body parts whose distance from the estimated ground is no greater than the corresponding contact threshold, and apply a constraint that the velocity in the contact normal direction is zero to the contact parts. ,in Let p be the three-dimensional velocity vector of the p-th contact point. To estimate the unit vector of the ground normal direction. Contact thresholds are set according to body part: 3 cm for the sole of the foot, 5 cm for the knee, 6 cm for the hip, 8 cm for the back, 6 cm for the shoulder, and 3 cm for the palm.

[0075]

[0076] Where w1=1.0, w2=0.8, w3=1.0, and w4=0.5.

[0077] like Figure 7 As shown, the motion field-dynamics cross-validation loss Its function is to constrain the consistency between the inter-frame velocity of the sub-pose and the velocity derived from the blur kernel, forming a third constraint independent of the optical and mechanical domains.

[0078] For each body part subregion i, calculate two velocity quantities: (a) Dynamically derived velocity —From the current iteration of the sub-pose sequence, perform finite difference analysis on the 3D joint displacements between adjacent subframes of the i-th body part sub-region to obtain the 3D velocity vector, then project it onto the 2D image plane to obtain the projection velocity direction. And projection speed (b) Fuzzy kernel derivation speed —The fuzzy kernel orientation angle of the sub-region is directly obtained from the output of the fuzzy kernel motion field analysis module. and fuzzy length Divide by exposure time The obtained speed and direction .

[0079] The calculation of the cross-validation loss between the sports field and dynamics is as follows:

[0080]

[0081] in The weight of the i-th body part sub-region, high-speed part ( The weight of the large (high) part is greater than the weight of the low-speed part, specifically taking... ; The weight for directional consistency is set to 1.0. The weight for consistency in speed magnitude is set to 0.5. The weight of sub-regions marked as low confidence is reduced to 0.2 times the original value.

[0082] The physical meaning of this loss is: if the sub-pose unfolding correctly reflects the actual motion during exposure, then the velocity between subframes should be consistent with the velocity encoded by the blur kernel. Three factors—optical appearance ( ), mechanical laws ( ), velocity field ( — This constitutes three independent constraint dimensions, which constrain the same set of sub-pose variables from three different angles: “similarity between the synthesized image and the actual image”, “motion satisfies Newtonian mechanics”, and “velocity is consistent with the blur kernel”. The intersection of the three sets of constraints shrinks the solution space from an infinite number of solutions for the optical constraints to a near-unique solution.

[0083] like Figure 8 As shown, after each iteration of optimization, the airborne or contact state of each subframe is re-determined based on the distance between each body part in the current sub-pose and the estimated ground. The determination method is as follows: check the distance from six types of contact candidate parts (both soles of the feet, both knees, center of the hip, center of the back, both shoulders, and both palms) to the estimated ground plane. If the distance from all candidate parts to the ground is greater than the corresponding contact threshold, the subframe is determined to be in an airborne state; if at least one candidate part is not greater than the corresponding contact threshold, the subframe is determined to be in a contact state, and the part whose distance is not greater than the threshold is marked as the current contact part.

[0084] The method for estimating the ground plane is as follows: The three-dimensional position sequence of key points on the human foot is detected in clear frames. The set of positions where the foot position is lowest during the standing and walking phases is taken, and the plane equation is fitted using the least squares method. For uneven ground scenes, local ground normal vectors are fitted using the foot positions from consecutive contact frames for adaptive correction.

[0085] Dynamic determination of the contact state allows the precise timing of take-off and landing to emerge automatically during the iteration process. A finite state machine with four states (contact state, take-off transition, airborne state, and landing transition) is constructed, and temporal logical consistency of state transitions is enforced on the subframe sequence—non-physical state jitter is not allowed within three consecutive subframes.

[0086] like Figure 7 As shown, after each round of iterative optimization, the dynamic derived velocity is calculated for each sub-region i of the body part. Compared with the current fuzzy kernel export speed The deviation between them. If any of the following conditions are met, the blur kernel correction for that sub-region is triggered:

[0087] Condition 1: Directional deviation (In this embodiment) (Take 20 degrees) This is the directional deviation threshold.

[0088] Condition 2: Speed ​​ratio deviation (In this embodiment) (Take 0.4, meaning the ratio of the dynamics-derived speed to the fuzzy kernel-derived speed deviates from 1 by more than 40%). This is the speed ratio deviation threshold.

[0089] When a correction is triggered, the fuzzy kernel parameters for that sub-region are updated using a weighted average:

[0090]

[0091]

[0092] Where η is the correction step size, and in this embodiment, η = 0.3 is used to ensure that the correction process converges smoothly rather than abruptly. Let be the blur length of the i-th body part sub-region. The corrected blur kernel parameters are then re-substituted into the motion field-dynamics cross-validation loss. The calculation is also used to update the motion field guidance information in the sub-pose initialization.

[0093] The closed loop of physical constraint feedback to correct the blur kernel parameters forms a complete closed loop of "image analysis → sub-pose unfolding → physical verification → velocity deviation detection → blur kernel correction → motion field update → sub-pose re-initialization → re-optimization". The technical significance of this closed loop is that the initial blur kernel estimation relies on image signal processing technology, and estimation errors are inevitable in severely blurred or low signal-to-noise ratio areas (such as small areas of the hand or toes with extremely high movement speed). If the blur kernel orientation angle deviates by 10 to 20 degrees, optical forward consistency constraints alone may find a local optimum in the wrong direction (because the sensitivity of the blurred image to direction is limited), but physical constraints will find that the mechanical trajectory in this direction is unreasonable (for example, the blur kernel moving horizontally in the take-off phase contradicts the vertical upward mechanical thrust), thus correcting the blur kernel orientation through velocity deviation signal feedback.

[0094] Iterative optimization uses the Adam optimizer, performing gradient descent on the total loss function with respect to all parameters of the N sub-poses. The number of iterations is set to 100 in this embodiment. A blur kernel correction check and contact state re-determination are performed every 10 iterations. The process terminates early when the total loss decreases by less than 0.1% over five consecutive iterations.

[0095] like Figure 9 As shown, the global temporal optimization module stitches the pose of clear frames and the sub-pose of blurred frames into a complete sequence and performs final optimization on a global scale.

[0096] Sequence stitching steps: Stitch the single pose of the clear frame with the N sub-poses of the blurred frame in chronological order. Perform local B-spline smoothing to eliminate speed jumps 3 to 5 sub-frames before and after the stitching point.

[0097] Global ground reaction force continuity verification steps: Calculate the global virtual ground reaction force time series for the complete stitched sequence. Apply a global continuity constraint—the time change rate of the ground reaction force does not exceed a preset force rate upper limit (in this embodiment, it is set to 50 times body weight per second, corresponding to the physiological upper limit of normal human muscle force rate). After low-pass filtering the ground reaction force time series, compare it with the original ground reaction force. Frames with large differences are marked as suspicious frames and sent back to the physical constraint sub-pose unfolding engine to increase the physical constraint weight for re-optimization.

[0098] State machine consistency verification steps: Construct a finite state machine containing four states: contact state, take-off transition, airborne state, and landing transition. Enforce temporal logical consistency of state transitions across the complete sequence. Perform a sliding window check on the global sequence. If a label sequence that violates the state machine transition rules appears within the window, correct the state label of the violating frame to the most recent valid state that conforms to the transition rules.

[0099] Output formatting steps: The globally optimized super-resolution pose sequence is downsampled back to the target frame rate, prioritizing the selection of the subframe with the highest physical plausibility score in each group as the representative pose. Complete subframe data is retained as an optional high frame rate output. Finally, the joint rotation parameters are converted to standard BVH (BioVision Hierarchy) format via inverse kinematics and output.

[0100] Example

[0101] The following example of a somersault illustrates the complete working process of the method of this invention. Figure 10 As shown.

[0102] Scene setup: A single cinematic camera (24 frames per second, 180-degree shutter angle, exposure time approximately 20 milliseconds) was used to film an actor performing a front flip in an indoor training area. The total duration of the front flip was approximately 1.0 second, corresponding to 24 frames of video.

[0103] Frame quality assessment results: Frames 1 to 7 and 23 to 24 are clear frames (running and stabilizing phases, M < 0.6); Frames 8 to 22 are blurry frames (high-speed action phases, M between 0.6 and 0.95), with frames 11 to 18 having the highest blur score (airborne flip phases).

[0104] Clear frame pose estimation: Full 3D pose is obtained in frames 1-7 and 23-24. Global morphological parameters: Height 175 cm, estimated weight 70 kg.

[0105] Motion field analysis of blurred frames: Taking frame 14 (in mid-air) as an example, the maximum blur length of the foot sub-region. =48 pixels (direction angle) =135° (corresponding to the direction of flipping motion), blur length of the torso sub-region. =12 pixels. Number of subframes N= =10, equivalent sampling rate is 240 frames / second. Subframe time interval =20 milliseconds / 10 = 2 milliseconds.

[0106] Acceleration accuracy analysis: At the original 24 frames / second, the typical foot acceleration is 100m / s². 2 The center difference cutoff error is approximately 35 m / s 2 (Relative error 35%), at this level of precision, the GRF calculated by inverse dynamics is unreliable on an order of magnitude—it cannot distinguish between "levitation (GRF should be zero)" and "light contact (GRF approximately 100 Newtons)". After the sub-attitude is expanded to N=10, the truncation error decreases to approximately 0.3 m / s². 2(Considering the optimized convergence accuracy, the effective relative error is approximately 4%). At this accuracy, GRF can accurately distinguish between airborne and contact states, and physical constraints can be effectively applied.

[0107] Triple joint iterative optimization process: Taking frame 14 as an example. Iteration begins after initializing 10 sub-poses:

[0108] Iterations 1-15: Optical Loss The image rapidly decreases in size, and the synthesized blurred image gradually approximates the actual blurred frame. Motion field-dynamics cross-validation loss. As synchronization decreases, the subframe velocity begins to approach the blur kernel velocity. Physical loss. The descent is slow (due to the poor physical plausibility of the initial sub-attitude).

[0109] In the 10th iteration (first fuzz kernel correction check): the dynamics-derived velocity direction of the foot sub-region is 140 degrees, deviating by 5 degrees from the initial fuzz kernel direction angle of 135 degrees (not exceeding the 20-degree threshold), and no correction is triggered. However, the dynamics-derived velocity direction of the left-hand sub-region is 85 degrees, deviating by 15 degrees from the initial estimate of 70 degrees, with a velocity ratio deviation of 50% (exceeding the 40% threshold), triggering correction: the left-hand fuzz kernel direction angle is corrected to 0.3×85+0.7×70=74.5 degrees, and the fuzz length is corrected with the same weight. The initial update of the left-hand sub-pose after correction, L... mf The loss in the left hand decreased significantly.

[0110] Iterations 15 to 50: Physical constraints gradually take effect. Frame 14 corresponds to the takeoff phase, where the projectile consistency constraint drives the root node trajectory to tend towards a parabola. middle The sub-items continue to decline. The triple loss converges in a coordinated manner.

[0111] Iterations 50 to 80: The total loss tends to plateau. The early termination condition is triggered on the 70th iteration.

[0112] Dynamic evolution of contact state: During the unfolding and optimization of frames 8 to 22, the moment of takeoff was precisely located to the 7th subframe of frame 10 (approximately 10 frames + 14 milliseconds in video time), and the moment of landing was precisely located to the 3rd subframe of frame 19 (approximately 19 frames + 6 milliseconds), with a positioning accuracy of approximately 4 milliseconds. During the roll, the contact area automatically migrates sequentially from the palm → shoulder → back → buttocks → sole of the foot, and the migration process is presented as a smooth transfer of contact area in the subframe sequence.

[0113] Global sequence optimization: The pose of clear frames (9 frames) and the sub-poses of blurred frames (15 frames × 10 sub-frames = 150 sub-poses) are stitched together to form a complete sequence containing 159 time points. Global GRF continuity check: An issue with excessive rate of change (exceeding the force rate limit) was found at the junction of frame 22 (end of landing buffer) and frame 23 (initial stabilization, clear frame). Frame 22 was re-optimized and the issue was resolved. State machine consistency check: No illegal transitions were found. The final output is downsampled to 24 frames / second.

[0114] Comparison of effects:

[0115] Option 1 (Direct Frame-by-Frame Inference): The failure rate of key point detection exceeds 85% from frame 8 to 22, and there are serious abrupt changes and physical violations in the output pose.

[0116] Scheme 2 (blurred frame unrolling with only optical constraints, i.e., D1 type scheme): The optical reconstruction quality is better, but there are multiple ambiguities in the airborne frames (the difference in PSNR (Peak Signal-to-Noise Ratio) between the two sets of different sub-pose sequences is less than 0.5dB, but the difference in the corresponding three-dimensional trajectory exceeds 10 cm), and some frames show feet going through the ground or unnatural joint over-limits.

[0117] Scheme 3 (optical constraints + physical constraints but no motion field cross-validation, i.e., simple D1+D2 series scheme): The physical rationality is significantly improved, but in sub-regions such as the hand where the initial fuzzy kernel estimation is inaccurate, the gradient direction conflict between the physical constraints and the optical constraints leads to non-convergence of local oscillations, and hand pose jitter artifacts appear.

[0118] Scheme 4 (this invention, triple constraints and physical feedback correction): The average joint positioning error across all blurred frames is controlled to be between 4 and 6 cm, the root mean square error of the root node trajectory deviating from the parabola during the take-off phase is less than 2 cm, and there are no physical violations throughout the sequence. In the hand region where local oscillations occurred in Scheme 3, the physical feedback correction mechanism automatically corrected the initial blur kernel direction deviation, the triple constraints converged smoothly, and there were no jitter artifacts in the hand posture.

[0119] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A markerless monocular video three-dimensional pose reconstruction method of motion capture, characterized in that, The method comprises the following steps: frame-level motion blur degree evaluation is performed on each frame image of the monocular video, and each frame is divided into a clear frame and a blurred frame according to the evaluation result; three-dimensional pose estimation is performed on the clear frame to obtain a three-dimensional pose corresponding to the clear frame and global human body shape parameters; regionally blurred kernel estimation is performed on the blurred frame, the human body region is divided into a plurality of sub-regions corresponding to body parts, blurred kernel parameters are independently estimated for each sub-region, a full-body motion field is constructed, and the number of sub-frames is dynamically determined according to the blurred kernel parameters; the three-dimensional pose of the clear frame is taken as a time sequence boundary condition, the blurred frame is subjected to sub-pose initialization in combination with the full-body motion field, and a plurality of sub-poses are generated; optical forward consistency constraint, dynamic physical rationality constraint and motion field-dynamics cross verification constraint are simultaneously applied to the plurality of sub-poses for triple joint iterative optimization, wherein the optical forward consistency constraint is an image consistency constraint between a simulated blurred image synthesized by the plurality of sub-poses and an actual blurred frame image, the dynamic physical rationality constraint is that inverse dynamics analysis is performed according to the global human body shape parameters and a motion sequence of the plurality of sub-poses to obtain virtual ground reaction force and apply a physical rationality constraint to the virtual ground reaction force, and the motion field-dynamics cross verification constraint is a consistency constraint between inter-frame velocities of the plurality of sub-poses and motion velocities of corresponding body parts in the full-body motion field; and a physically consistent sub-pose sequence is obtained through the triple joint iterative optimization; the three-dimensional pose of the clear frame and the sub-pose sequence obtained by unfolding the blurred frame are time sequence spliced, and a three-dimensional pose reconstruction result is output after global dynamic consistency optimization.

2. The markerless monocular video three-dimensional pose reconstruction method of claim 1, wherein, The blurred kernel parameters comprise a direction angle and a blur length; in the regionally blurred kernel estimation, the frequency domain zero stripe analysis method is used to estimate the blurred kernel parameters for a sub-region with an area greater than a preset area threshold, and the gradient direction histogram method is used to estimate the blurred kernel parameters for a sub-region with an area not greater than the preset area threshold; The dynamic determination of the subframe expansion quantity is as follows: wherein N is the subframe expansion quantity, is the maximum blur length in all sub-regions, is a preset maximum allowed joint displacement amount between subframes, is a ceiling function.

3. The markerless monocular video 3D pose reconstruction method of claim 1, wherein, The sub-pose initialization is motion field guided non-uniform sub-pose initialization, which comprises: according to a motion amplitude indicated by the blurred kernel parameters of each body part sub-region, a frame inter-displacement increment of a high-speed body part with a large motion amplitude is allocated to be greater than that of a low-speed body part with a small motion amplitude; in each round of iterative optimization, the distance between each body part and an estimated ground is dynamically determined to determine a current sub-frame in a flying state or a contact state, and a body part with a distance from the estimated ground not greater than a corresponding contact threshold is determined as a current contact body part.

4. The markerless monocular video 3D pose reconstruction method of claim 1, wherein, The implementation manner of the optical forward consistency constraint comprises: each sub-pose is input into a differentiable human body model renderer to obtain a two-dimensional projection image corresponding to each sub-pose; a plurality of two-dimensional projection images are superimposed to generate a synthesized blurred image according to uniform weights; and pixel level difference and edge blur mode difference between the synthesized blurred image and the actual blurred frame image are calculated as a loss value of the optical forward consistency constraint.

5. The markerless monocular video 3D pose reconstruction method of claim 1, wherein, The inverse dynamics analysis includes: determining the mass and moment of inertia parameters of each limb based on global human morphology parameters; using the shortened inter-frame time interval after subframe expansion to perform second-order finite difference calculations on the motion sequences of multiple sub-poses to obtain joint accelerations; and recursively calculating virtual ground reaction forces through Newton-Euler inverse dynamics. The physical rationality constraints include projectile consistency constraints during the takeoff phase, friction cone constraints during the contact phase, non-negative contact force constraints, and multi-part rolling contact constraints. Among them, the projectile consistency constraint during the takeoff phase requires that the horizontal acceleration of the root node in the subframe of the takeoff state approaches zero and the vertical acceleration approaches the negative value of the gravitational acceleration. The friction cone constraint during the contact phase requires that the tangential component of the virtual ground reaction force in the subframe of the contact state does not exceed the product of the normal component and the friction coefficient. The non-negative contact force constraint requires that the normal component of the virtual ground reaction force in the subframe of the contact state is non-negative. The multi-part rolling contact constraint applies a constraint that the velocity in the contact normal direction of body parts not greater than the estimated ground is close to zero. The body parts include the soles of the feet, knees, hips, back, shoulders, and hands.

6. The markerless monocular video 3D pose reconstruction method of claim 1, wherein, The triple joint iterative optimization also includes a physical constraint feedback correction step: after each round of iterative optimization, for each body part sub-region, the projection direction and projection length of the joint displacement vector between adjacent sub-poses in the sub-region on the two-dimensional image plane are compared with the fuzzy kernel parameters of the sub-region. When the deviation between the dynamic derived velocity and the fuzzy kernel derived velocity exceeds a preset velocity deviation threshold, the fuzzy kernel parameters of the sub-region are corrected to the weighted average of the dynamic derived value and the original fuzzy kernel parameter value, and the corresponding sub-pose is re-initialized according to the corrected fuzzy kernel parameters.

7. The markerless monocular video 3D pose reconstruction method of claim 1, wherein, The loss calculation method of the motion field-dynamics cross-validation constraint is as follows: For each body part sub-region, calculate the mean square error between the projection velocity of the sub-pose inter-frame velocity vector in the sub-region on the two-dimensional image plane and the motion velocity of the corresponding body part in the whole body motion field. The mean square error of all sub-regions is weighted and summed as the loss value, where the weight of high-speed body parts is greater than the weight of low-speed body parts. The global dynamic consistency optimization includes: calculating the global virtual ground reaction time series for the spliced ​​complete attitude sequence, applying ground reaction time continuity constraints to ensure that the ground reaction change rate does not exceed the preset force rate upper limit, and constructing a finite state machine containing four states: contact state, take-off transition, airborne state, and landing transition, and forcing temporal logic consistency of state transitions on the complete sequence.

8. A markerless monocular video three-dimensional pose reconstruction system for motion capture, the system comprising: include: The frame quality assessment module is used to assess the motion blur level of each frame of a monocular video and classify each frame into clear frames and blurred frames based on the assessment results. The basic pose estimation module is used to perform 3D pose estimation on clear frames and obtain the 3D pose and global human body shape parameters corresponding to the clear frames. The fuzzy kernel motion field analysis module is used to perform regional fuzzy kernel estimation on fuzzy frames. It divides the human body region into multiple sub-regions corresponding to body parts and independently estimates the fuzzy kernel parameters for each sub-region to construct the whole-body motion field. It also dynamically determines the number of sub-frames to be unfolded based on the fuzzy kernel parameters. The physical constraint sub-pose unfolding engine uses the 3D pose of a clear frame as a temporal boundary condition, combined with the whole-body motion field to initialize and generate multiple sub-poses. It simultaneously applies optical forward consistency constraints, dynamic physical rationality constraints, and motion field-dynamic cross-validation constraints to multiple sub-poses for triple joint iterative optimization to obtain a physically consistent sub-pose sequence. The optical forward consistency constraint is the image consistency constraint between the simulated blurred image synthesized from multiple sub-poses and the actual blurred frame image. The dynamic physical rationality constraint is the physical rationality constraint applied to the virtual ground reaction force obtained by inverse dynamics analysis based on global human body morphology parameters and the motion sequence of multiple sub-poses. The motion field-dynamic cross-validation constraint is the consistency constraint between the inter-frame velocity of multiple sub-poses and the motion velocity of the corresponding body parts in the whole-body motion field. The global temporal optimization module is used to temporally stitch together the 3D pose of the clear frame and the sub-pose sequence obtained by unfolding the blurred frame, and output the 3D pose reconstruction result after global dynamic consistency optimization.

9. The markerless monocular video three-dimensional pose reconstruction system of claim 8, wherein, The physical constraint sub-pose unfolding engine includes: a sub-pose initialization unit, used to perform non-uniform interpolation initialization by proportionally allocating inter-frame displacement increments based on the motion amplitude indicated by the fuzzy kernel parameters of each body part in conjunction with the whole-body motion field; a triple constraint iterative optimization unit, used to calculate optical forward consistency loss, dynamic physical rationality loss, and motion field-dynamic cross-validation loss, and use the weighted sum of the three as the total loss function for gradient backpropagation optimization; a contact state dynamic determination unit, used to determine the airborne or contact state of each sub-frame and identify the contacting body part based on the distance between each body part and the estimated ground; and a fuzzy kernel feedback correction unit, used to correct the fuzzy kernel parameters and reinitialize the corresponding sub-pose when the deviation between the dynamic derived velocity and the fuzzy kernel derived velocity exceeds a preset velocity deviation threshold.

10. The markerless monocular video three-dimensional pose reconstruction system of claim 8, wherein, The global timing optimization module includes: a sequence splicing submodule, used to splice the clear frame attitude and the sub-attitude of the blurred frame unfolded in time order and perform local smoothing at the splicing point; a global reaction force continuity verification submodule, used to calculate the virtual ground reaction force time series of the complete sequence and apply a force rate upper limit constraint; and a state machine consistency submodule, used to construct a finite state machine containing four states: contact state, take-off transition, airborne state, and landing transition, and eliminate inter-frame jitter of state labels.