Motion capture optimization method and model training method based on monocular video
By employing a two-stage optimization method, the foot trajectory in monocular video is optimized using a ground contact detection model and inverse dynamics algorithm. This solves the problems of foot clipping and jitter in monocular video-based motion capture technology, and improves the accuracy and robustness of 3D motion data.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING UNIV OF POSTS & TELECOMM
- Filing Date
- 2025-01-15
- Publication Date
- 2026-05-12
AI Technical Summary
基于单目视频的动作捕捉技术中,人像脚部在三维重构过程中容易出现穿模和抖动现象,导致三维动作数据的准确性降低。
A two-stage optimization method is adopted. First, the ground contact probability is obtained and the foot trajectory is optimized through the ground contact detection model. Then, the three-dimensional motion data is reconstructed through the inverse dynamics algorithm, and the jitter is eliminated by combining forward dynamics and filtering algorithms to optimize the foot position.
It improves the accuracy of 3D motion data, solves the problems of foot trajectory jumps and jitters, and ensures robustness in different scenarios.
Smart Images

Figure CN119992651B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a motion capture optimization method and model training method based on monocular video. Background Technology
[0002] Motion capture technology is an interdisciplinary technology that integrates image processing, computer vision, and human kinematics. Its core lies in capturing human motion information, providing a data foundation for many fields such as virtual reality, augmented reality, animation production, and medical rehabilitation.
[0003] Among numerous motion capture technologies, monocular video-based motion capture technology has gradually become a research focus due to its economy, convenience, and applicability. This technology only requires a single camera or a single video stream as input to extract key points of a human figure, and then reconstructs the three-dimensional motion data of the figure through inverse kinematics (IK) algorithms.
[0004] However, due to the lack of depth information, motion capture technology based on monocular video often encounters problems such as decreased spatial positioning accuracy of key points and occlusion. At the same time, changes in environmental factors such as lighting conditions may also interfere with the extraction of key points. These problems can lead to clipping and jitter in the 3D reconstruction of human figures' feet, which in turn reduces the accuracy of the 3D motion data reconstructed by inverse kinematics algorithms. Summary of the Invention
[0005] This application provides a motion capture optimization method and model training method based on monocular video to solve the problem that clipping and shaking may occur in the 3D reconstruction of human feet, which leads to a decrease in the accuracy of 3D motion data reconstructed by the inverse dynamics algorithm.
[0006] The first aspect of this application provides a motion capture optimization method based on monocular video, the method comprising:
[0007] Acquire monocular video and input it into a pre-set ground touch detection model to obtain the ground touch detection result output by the ground touch detection model. The ground touch detection result is used to indicate the probability of the human figure's foot touching the ground in the monocular video. The ground touch detection model is trained using a labeled ground touch dataset, which is used to indicate whether the human figure's foot touches the ground in the training frame.
[0008] Based on the monocular video, the first three-dimensional motion data of the human figure is obtained, and based on the first three-dimensional motion data, the trajectory of the human figure's feet is obtained.
[0009] Based on the ground contact probability, the foot trajectory is optimized for ground contact, and based on the ground contact probability, the foot trajectory after ground contact optimization is optimized for shaking.
[0010] Based on the foot trajectory optimized by shaking, the first three-dimensional motion data is reconstructed using the inverse dynamics algorithm to obtain the second three-dimensional motion data of the human figure.
[0011] In one possible design, the foot trajectory is optimized for ground contact based on the probability of ground contact, including:
[0012] Based on the foot trajectory, the position of the first foot is obtained through a forward dynamics algorithm, and based on the ground contact probability and the position of the first foot, a clustering algorithm is used to construct a predicted ground surface.
[0013] The foot trajectory is optimized for ground contact based on the ground contact probability, the first foot position, the estimated ground, and the foot speed indicated by the foot trajectory.
[0014] In one possible design, the first foot position includes the toe position and the heel position, then the ground contact probability includes the toe contact probability and the heel contact probability, and the foot velocity includes the toe velocity and the heel velocity.
[0015] Based on the ground contact probability, the first foot position, the estimated ground surface, and the foot speed indicated by the foot trajectory, the foot trajectory is optimized for ground contact, including:
[0016] The sliding loss is calculated based on the probability of toe contact with the ground, the probability of heel contact with the ground, the toe speed, and the heel speed.
[0017] The foot is gridded based on the position of the toe and the heel, and the clipping loss is obtained based on the estimated ground and the gridded foot.
[0018] Based on the sliding loss, clipping loss, and preset posture loss, the weighted result of each weight combination is obtained by weighted calculation through multiple preset weight combinations;
[0019] The foot trajectory is optimized for ground contact based on the smallest weighted result among multiple weighted results.
[0020] In one possible design, the foot trajectory after ground contact optimization is subjected to shaking optimization based on the ground contact probability, including:
[0021] Based on the foot trajectory optimized by ground contact, the position of the second foot is obtained through a forward dynamics algorithm, and the end jitter of the second foot position is eliminated through a first filtering algorithm to obtain the end position of the foot.
[0022] The foot state is obtained based on the ground contact probability and the position of the foot tip; the foot state is used to indicate the ground contact state of the human figure's two feet.
[0023] Based on the foot condition, the foot trajectory after ground contact optimization is subjected to vibration optimization.
[0024] In one possible design, the foot position can be one of two feet fully on the ground, two feet completely off the ground, or partially on the ground;
[0025] When both feet are fully on the ground, the foot trajectory, optimized for ground contact, is subjected to shaking optimization based on the foot position, including:
[0026] Eliminate foot movement;
[0027] When both feet are fully off the ground, the foot trajectory, optimized for ground contact, is subjected to shaking optimization based on the foot position, including:
[0028] By using trajectory acceleration and acceleration constraints, the foot trajectory after ground contact optimization is subjected to jitter optimization;
[0029] When the foot is in partial contact with the ground, the foot trajectory after ground contact optimization is subjected to shaking optimization based on the foot state, including:
[0030] The foot trajectory after ground contact optimization is jitter optimized by using a second filtering algorithm and a frame interpolation algorithm.
[0031] In one possible design, based on monocular video, the first three-dimensional motion data of the human figure is obtained, including:
[0032] Based on the monocular video, the trajectory of the first human figure is obtained through pose estimation, as well as the two-dimensional loss, smoothing function and initial pose of the human figure;
[0033] Based on the first portrait trajectory, two-dimensional loss, and smoothing function, the rotation and displacement of the camera parameters are optimized to obtain the second portrait trajectory.
[0034] Based on the two-dimensional loss, smoothing function, and initial pose, the rotation, displacement, and pose of the second human figure trajectory are optimized to obtain the first three-dimensional motion data.
[0035] A second aspect of this application provides a model training method, the method comprising:
[0036] Obtain the labeled Touch Touch dataset; the Touch Touch dataset includes multiple training frames, each of which includes a human image;
[0037] A ground touch detection model is trained based on the labeled ground touch dataset; the ground touch detection model is used in any of the motion capture optimization methods based on monocular videos in the first aspect.
[0038] In one possible design, a ground contact detection model is trained based on the labeled ground contact dataset, including:
[0039] The labeled ground-touch dataset is rendered from multiple perspectives to obtain a keypoint model; the keypoint model is used to indicate the rotation and trajectory information of the key points of the human figure.
[0040] Based on the keypoint model, intra-frame and inter-frame joint interaction information is obtained through a pre-built DSTformer network and a transformer attention mechanism.
[0041] Based on joint interaction information, a ground contact detection model is trained using a pre-defined binary cross-entropy loss function.
[0042] In one possible design, the method further includes the following during each iteration of training:
[0043] Calculate the first loss function of the ground contact probability output by the ground contact detection model, and the velocity loss function between frames;
[0044] The second loss function is calculated by weighting the first loss function and the velocity loss function; wherein the weight of the first loss function is greater than the weight of the velocity loss function.
[0045] Use the second loss function as the new first loss function.
[0046] In one possible design, the labeled touchdown dataset is obtained, including:
[0047] Obtain the ground contact dataset;
[0048] Based on the ground contact dataset, a spring model of the human figure is constructed to obtain the force distribution on the feet of the human figure.
[0049] Based on the force applied to the feet, labeled data is obtained for each training frame; the labeled data is used to indicate whether the feet of the human figure in the corresponding training frame are touching the ground.
[0050] The ground contact dataset is labeled based on the labeled data of each training frame.
[0051] A third aspect of this application provides a motion capture optimization device based on monocular video, the device comprising:
[0052] The result determination module is used to acquire monocular video and input the monocular video into a preset ground touch detection model to obtain the ground touch detection result output by the ground touch detection model. The ground touch detection result is used to indicate the probability of the human figure's foot touching the ground in the monocular video. The ground touch detection model is trained by an annotated ground touch dataset, which is used to indicate whether the human figure's foot touches the ground in the training frame.
[0053] The trajectory determination module is used to obtain the first three-dimensional motion data of the human figure based on the monocular video, and to obtain the foot trajectory of the human figure based on the first three-dimensional motion data.
[0054] The foot optimization module is used to optimize the foot trajectory based on the ground contact probability, and to optimize the jitter of the optimized foot trajectory based on the ground contact probability.
[0055] The data reconstruction module is used to reconstruct the first three-dimensional motion data of the human figure based on the foot trajectory optimized by shaking, and obtain the second three-dimensional motion data of the human figure.
[0056] A fourth aspect of this application provides a model training apparatus, the apparatus comprising:
[0057] The data acquisition module is used to acquire the labeled Touch Touch dataset; the Touch Touch dataset includes multiple training frames, each of which includes a human image;
[0058] The model training module is used to train a ground contact detection model based on the labeled ground contact dataset; the ground contact detection model is used in the third aspect of the motion capture optimization device based on monocular video.
[0059] A fifth aspect of this application provides an electronic device, including: a memory, and a memory communicatively connected to a processor;
[0060] The memory stores instructions that the computer executes;
[0061] When the processor executes computer execution instructions stored in memory, it is used to implement either the motion capture optimization method based on monocular video in the first aspect, or the model training method in the second aspect.
[0062] The sixth aspect of this application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the motion capture optimization method based on monocular video according to any one of the first aspects, or the model training method according to any one of the second aspects.
[0063] The seventh aspect of this application provides a computer program product, including a computer program, which, when executed by a processor, is used to implement the motion capture optimization method based on monocular video of any one of the first aspects, or the model training method of any one of the second aspects.
[0064] This application provides a motion capture optimization method and model training method based on monocular video. The method includes: inputting monocular video into a ground-touch detection model to obtain ground-touch detection results indicating the ground-touch probability of a person's feet; obtaining first three-dimensional motion data of the person based on the monocular video, and then obtaining the foot trajectory of the person; optimizing the foot trajectory based on the ground-touch probability, and then optimizing the jitter of the optimized foot trajectory; reconstructing the first three-dimensional motion data of the person using an inverse dynamics algorithm based on the jitter-optimized foot trajectory to obtain second three-dimensional motion data of the person. The method achieves the following technical effects: It performs two-stage optimization of the foot trajectory; the first stage, ground-touch optimization, solves the problems of foot trajectory jumps and clipping; the second stage, jitter optimization, solves the problem of jitter at the end of the foot trajectory; it reconstructs three-dimensional motion data using an inverse dynamics algorithm based on the two-stage optimized foot trajectory, improving the accuracy of the three-dimensional motion data; and it outputs the ground-touch detection results of the monocular video from a ground-touch detection model trained on an annotated ground-touch dataset, solving the problem of difficulty in determining the boundary frames of the foot leaving the ground. Attached Figure Description
[0065] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0066] Figure 1 This is a schematic diagram of a scenario provided for an embodiment of this application;
[0067] Figure 2 A flowchart illustrating the motion capture optimization method based on monocular video provided in this application embodiment. Figure 1 ;
[0068] Figure 3 A flowchart illustrating the motion capture optimization method based on monocular video provided in this application embodiment. Figure 2 ;
[0069] Figure 4 A schematic diagram illustrating the principle of ground contact optimization provided in this application embodiment;
[0070] Figure 5 A schematic diagram illustrating the principle of jitter optimization provided in this application embodiment;
[0071] Figure 6 A flowchart illustrating the model training method provided in the embodiments of this application. Figure 1 ;
[0072] Figure 7A flowchart illustrating the model training method provided in the embodiments of this application. Figure 2 ;
[0073] Figure 8 A schematic diagram of the process for reconstructing the second data of three-dimensional motion provided in an embodiment of this application;
[0074] Figure 9 A schematic diagram illustrating the process of monocular video annotation provided in an embodiment of this application;
[0075] Figure 10 A schematic diagram of the structure of the motion capture optimization device based on monocular video provided in the embodiments of this application;
[0076] Figure 11 This is a schematic diagram of the structure of the model training device provided in the embodiments of this application;
[0077] Figure 12 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0078] Figure label:
[0079] 110 - Monocular camera; 120 - First data processing server; 130 - Second data processing server;
[0080] 101 - Result Determination Module; 102 - Trajectory Determination Module; 103 - Foot Optimization Module; 104 - Data Reconstruction Module;
[0081] 111 - Data Acquisition Module; 112 - Model Training Module;
[0082] 121-Processor; 122-Memory; 123-Communication components; 124-Bus. Detailed Implementation
[0083] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0084] In this application, the terms "first" and "second" are used to distinguish identical or similar items with substantially the same function and effect. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, nor do they necessarily imply difference. It should be noted that in this application, the terms "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design scheme described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner. In this application, "at least one" means one or more, and "more than one" means two or more.
[0085] It should be noted that the phrase "at the moment" in this application can refer to the instant a certain situation occurs, or to a period of time after the occurrence of a certain situation; this application does not impose a specific limitation on this. Furthermore, the motion capture optimization method and model training method based on monocular video provided in this application are merely examples, and the motion capture optimization method and model training method based on monocular video may also include more or less content. The user information (including but not limited to user device information and user personal information) and data (including but not limited to data used for analysis, stored data, and displayed data) involved in one or more embodiments of this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with relevant laws, regulations, and standards, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0086] To facilitate a clear description of the technical solution of this application, some of the terms and technologies involved in this application are briefly introduced below:
[0087] Motion capture technology refers to the technology that, by combining deep learning models and subsequent optimization methods, can extract three-dimensional (3D) human motion from a video that matches the original video.
[0088] Skinned Multi-Person Linear Model (SMPL): This refers to a 3D mesh model used to represent human posture and shape. SMPL defines a low-dimensional parameter space, including shape parameters and posture parameters. Shape parameters describe the overall shape of the human body, such as height, weight, and body size; posture parameters describe the human body's posture, i.e., the rotation angles of the joints.
[0089] Ground contact detection refers to the process of training a neural network to identify the ground contact status of a person's toes and heels.
[0090] Inverse Kinematics (IK) algorithm: This refers to an algorithm that first determines the position of a sub-bone, and then reversely calculates the position of its n-level parent bone in the bone chain, thereby determining the entire bone chain. In this application, the IK algorithm is mainly applied to foot pose adjustment.
[0091] Dual-stream Spatio-temporal Transformer (DSTformer) network: refers to a two-dimensional lifting network with a time-series transformer structure that can capture the temporal characteristics of human motion.
[0092] 2D-lifting network: refers to a 3D keypoint estimation network that learns prior knowledge about the human body by training the mapping from 2D keypoints to 3D keypoints.
[0093] The technical solutions of this application will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. This application will now be described with reference to the accompanying drawings.
[0094] To clearly understand the technical solution of this application, the solutions of the prior art will be described in detail first.
[0095] Motion capture technology is an interdisciplinary field that integrates image processing, computer vision, and human kinematics. Its core lies in capturing human motion information, providing a data foundation for numerous fields such as virtual reality, augmented reality, animation, and medical rehabilitation. Traditional motion capture technologies primarily rely on optical or inertial motion capture equipment. Optical motion capture equipment typically requires multiple high-speed cameras, marker points, and complex computing devices; it offers high accuracy but is also costly and requires specialized facilities and environmental conditions. Inertial motion capture equipment relies on multiple sensors attached to the human body, measuring acceleration and angular velocity to reconstruct human movement; however, it is susceptible to sensor errors and magnetic field interference.
[0096] Among numerous motion capture technologies, monocular video-based motion capture technology has gradually become a research focus due to its economy, convenience, and applicability. This technology only requires a single camera or a single video stream as input to extract key points such as joints and skeleton of a human figure, and then reconstructs the three-dimensional motion data of the human figure through inverse kinematics (IK) algorithms.
[0097] However, due to the lack of depth information, motion capture technology based on monocular video often encounters problems such as decreased spatial positioning accuracy of key points and occlusion; at the same time, changes in environmental factors such as lighting conditions may also interfere with the extraction of key points.
[0098] Therefore, existing technologies may exhibit clipping and jitter issues during 3D reconstruction of human figures' feet, leading to reduced accuracy of the 3D motion data reconstructed using inverse kinematics (IK) algorithms. To address this problem, the research found that current solutions can be divided into two categories. The first category uses a ground contact detection model to obtain ground contact results and then uses the IK algorithm to fix the foot position based on these results. The second category uses incidental results such as foot velocity output by the model as supervisory information to optimize the foot trajectory.
[0099] The first problem with the first type of approach is the lack of publicly available labeled datasets, and the difficulty in determining the boundary frames where the feet are off the ground during data annotation. Specifically, the data is manually labeled, and annotating a short video of about 200 frames takes approximately 30 minutes, resulting in excessively high manual annotation costs. Furthermore, the distinction between the front and back feet of a person in the image is not obvious, especially in scenes where the feet are not clearly off the ground, making it difficult to determine whether they are touching the ground. This ambiguity further increases the annotation cost. Based on this, this application proposes to automate the annotation of the ground-touching dataset using a spring model. By establishing a spring model on the light-captured human body data, ground-touching labels are obtained, and an animation engine is used to perform multi-view rendering on the light-captured human body data to increase the richness of the dataset.
[0100] The second problem with the first type of approach is that it is difficult to optimize for abrupt changes in foot pose when applying the IK algorithm. Specifically, existing technologies obtain the boundary frames of foot liftoff, i.e., ground-touch labels, by thresholding the speed and height of the foot. The model only learns the mapping from keypoint coordinates to ground-touch labels. Once the action type is not covered by the dataset, it cannot perform the mapping, resulting in unreliable label quality and poor generalization of ground-touch optimization for various videos. Based on this, this application proposes to fine-tune a 2D-lifting network for the ground-touch task to ensure high accuracy in various wild scenarios.
[0101] The problem with the second type of approach is that most of these methods optimize foot trajectory by designing foot speed and height losses. However, due to the complexity of the motion, optimizing only soft constraints may lead to insufficient generalization and a strong dependence on ground contact detection results. Errors in a few ground contact detection results can affect the smoothness and naturalness of the entire trajectory. Based on this, this application proposes to optimize foot trajectory using a two-stage ground contact optimization method that combines loss optimization and IK (Independent Kinetic Kinematics). The first stage focuses on correcting the height of the foot trajectory and the height of the foot, while the second stage focuses on smoothing and fixing the foot position.
[0102] Based on the aforementioned inventive findings, this application presents its technical solution. To address the aforementioned problems, this application proposes a motion capture optimization method and model training method based on monocular video. This application designs a two-stage optimization process, which can optimize and obtain smooth and physically plausible 3D motion data of a human figure using only the RGB video sequence of monocular video and the ground touch detection results output by the ground touch detection model. Furthermore, this application also proposes an automated annotation method for ground touch datasets, which can generate labels and multi-view video sequences based on the ground touch dataset, increasing the richness of the dataset.
[0103] The following section introduces the application scenarios of the motion capture optimization method and model training method based on monocular video provided in this application.
[0104] Figure 1 This is a schematic diagram of a scenario provided for an embodiment of this application. It should be noted that... Figure 1 The examples shown are merely examples of scenarios in which this application can be applied, to help those skilled in the art understand the technical content of this application, but do not mean that this application cannot be used in other devices, systems, environments or scenarios.
[0105] like Figure 1 As shown, the application scenario includes: a monocular camera 110, a first data processing server 120, and a second data processing server 130.
[0106] The monocular camera 110 is used to capture motion video of motion capture actors and outputs the captured RGB video, i.e., monocular video. The monocular camera 110 has high resolution and frame rate, which can clearly record every subtle movement of the motion capture actors.
[0107] The first data processing server 120 is equipped with powerful computing capabilities and ample storage space, enabling it to process and analyze monocular video from the monocular camera 110. The first data processing server 120 is communicatively connected to the monocular camera 110 to execute a motion capture optimization method based on monocular video, obtaining second data of the human figure's three-dimensional motion.
[0108] The second data processing server 130 is also equipped with powerful computing capabilities and ample storage space. The second data processing server 130 is used to execute model training methods, can process large-scale data annotations, and train a high-precision ground contact detection model.
[0109] The first data processing server 120 and the second data processing server 130 can be different servers, physically separated, each with its own independent processor and storage device. Therefore, the first data processing server 120 and the second data processing server 130 may be located in the same data center or distributed in different geographical locations. After the second data processing server 130 completes the ground contact detection model training, it can transfer the data to the first data processing server 120 via file transfer protocols such as FTP (File Transfer Protocol) or SFTP (Secure File Transfer Protocol), or cloud storage services. Alternatively, the second data processing server 130 can be deployed as a server providing a Representational State Transfer API (RESTful API) or Remote Procedure Call (RPC) service. The first data processing server 120 sends video frames or feature data to the second data processing server 130 by sending Hypertext Transfer Protocol (HTTP) requests or RPC calls, and receives responses from the ground contact detection results.
[0110] The first data processing server 120 and the second data processing server 130 can also be different logical partitions, containers, or virtual machines on the same physical server. They share the same physical hardware resources but are logically isolated through virtualization technology. After the second data processing server 130 completes the grounding detection model training, the model can be directly stored in the server's Network File System (NFS) mount point or shared storage area such as Storage Area Network (SAN) / Network Attached Storage (NAS) storage, which can be accessed by both servers as needed; or, the first data processing server 120 and the second data processing server 130 can directly exchange grounding detection results through inter-process communication mechanisms such as shared memory, pipes, or sockets.
[0111] Figure 2 A flowchart illustrating the motion capture optimization method based on monocular video provided in this application embodiment. Figure 1 .like Figure 2As shown in the embodiments of this application, the executing entity can be a motion capture optimization device based on monocular video. This device can be located in an electronic device. Figure 1 The first data processing server 120 in the application. The motion capture optimization method based on monocular video provided in this embodiment includes the following steps:
[0112] S201. Acquire monocular video and input the monocular video into a preset ground contact detection model to obtain the ground contact detection result output by the ground contact detection model.
[0113] Specifically, the touchdown detection model is trained using a labeled touchdown dataset, which indicates whether the feet of the human figures in the training frames touch the ground. Through training on this dataset, the model learns a method for annotating the probability of foot touchdowns. Then, the monocular video is input into the model, which annotates the probability of foot touchdowns in each frame and outputs the detection result. This result indicates the probability of foot touchdowns in each frame of the monocular video.
[0114] S202. Based on the monocular video, obtain the first three-dimensional motion data of the human figure, and based on the first three-dimensional motion data, obtain the foot trajectory of the human figure.
[0115] Specifically, while acquiring the ground contact detection results, the first data processing server also extracts the first three-dimensional motion data of the human figure from the monocular video (the first three-dimensional motion data refers to a type of SMPL data), and then calculates the trajectory of the human figure's feet. Among them, the first three-dimensional motion data refers to the data set describing the human figure's motion state in three-dimensional space, which usually includes information such as the three-dimensional coordinates of the human figure's joints, motion velocity, and acceleration, and can comprehensively and accurately reflect the motion characteristics of the human figure.
[0116] It should be noted that the execution order between S201 and S202 is not limited. It can be as follows: Figure 2 The execution order shown is S201 first, then S202. However, other execution orders are also possible, such as S202 first, then S201, or S201 and S202 being executed simultaneously.
[0117] S203. Based on the ground contact probability, optimize the foot trajectory for ground contact, and based on the ground contact probability, optimize the shaking of the optimized foot trajectory for ground contact.
[0118] Specifically, the first-stage ground-touch optimization aims to address the jump and clipping issues in the subsequent IK algorithm. While the IK algorithm can change the position from the foot to the knee, it cannot predict how smoothly a sequence will adjust. Even with smoothing through filtering or frame interpolation, robustness across different scenarios cannot be guaranteed. Therefore, ground-touch optimization is performed on the foot trajectory based on the ground-touch probability to resolve jump and clipping issues.
[0119] The second stage of jitter optimization first requires calculating the foot position using the Forward Kinematics (FK) algorithm, and then using filtering to eliminate end-foot jitter. FK is a physics-based simulation method that predicts the trajectory of a person's feet based on initial conditions such as position, velocity, and acceleration, as well as forces such as gravity and muscle force, combined with information such as joint angles and bone length.
[0120] S204. Based on the foot trajectory optimized by shaking, the first three-dimensional motion data is reconstructed using the inverse dynamics algorithm to obtain the second three-dimensional motion data of the human figure.
[0121] Specifically, based on the foot trajectory optimized by shaking, the root node trajectory can be adjusted using the IK algorithm. Then, the first 3D motion data is reconstructed based on the adjusted root node trajectory, resulting in the final output second 3D motion data of the human figure. Here, the root node refers to the node at the tail of the human figure's spine, and its trajectory directly determines the human figure's trajectory. The human figure's trajectory refers to the movement of the human figure in two-dimensional or three-dimensional space; it can be considered a special type of 3D motion data.
[0122] This application provides a motion capture optimization method based on monocular video. The method includes: inputting monocular video into a ground-touch detection model to obtain ground-touch detection results indicating the ground-touch probability of a person's feet; obtaining first three-dimensional motion data of the person based on the monocular video, and then obtaining the foot trajectory of the person; optimizing the foot trajectory based on the ground-touch probability, and optimizing the jitter of the optimized foot trajectory; reconstructing the first three-dimensional motion data of the person using an inverse dynamics algorithm based on the jitter-optimized foot trajectory to obtain second three-dimensional motion data of the person. This achieves the following technical effects: The foot trajectory is optimized in two stages; the first stage, ground-touch optimization, solves the problems of foot trajectory jumps and clipping; the second stage, jitter optimization, solves the problem of jitter at the end of the foot trajectory; the three-dimensional motion data is reconstructed using an inverse dynamics algorithm based on the two-stage optimized foot trajectory, improving the accuracy of the three-dimensional motion data; and the ground-touch detection model trained on the labeled ground-touch dataset outputs the ground-touch detection results of the monocular video, solving the problem of difficulty in determining the boundary frames of the foot leaving the ground.
[0123] Figure 3 A flowchart illustrating the motion capture optimization method based on monocular video provided in this application embodiment. Figure 2 ,like Figure 3 As shown, the motion capture optimization method based on monocular video provided in this application embodiment is... Figure 2 This embodiment further refines the motion capture optimization method based on monocular video provided in the previous embodiment. The motion capture optimization method based on monocular video provided in this application includes the following steps.
[0124] S301. Acquire monocular video and input the monocular video into a preset ground contact detection model to obtain the ground contact detection result output by the ground contact detection model.
[0125] S302. Based on the monocular video, obtain the trajectory of the first human figure through pose estimation, as well as the two-dimensional loss, smoothing function, and initial pose of the human figure.
[0126] Specifically, the first 3D motion data can be obtained through pre-optimization of the first human image trajectory. Pre-optimization includes two steps: camera parameter optimization and human image trajectory optimization. Based on the monocular video, the first human image trajectory, 2D loss, smoothing function, and initial pose are obtained.
[0127] Two-dimensional loss (2D_loss) refers to the loss function calculated in two-dimensional space.
[0128] Smoothing functions include smoothing depth (smoothDepth), smoothing rotation (smoothRh), smoothing displacement (smoothTh), and smoothing poses (smoothposes). smoothDepth smooths depth data to reduce noise or irregularities, making it smoother or more continuous. smoothRh smooths rotation (Rh) to reduce abrupt changes or jitter in rotational shifts, making movement more natural. smoothTh smooths displacement (Th) to reduce jumps or instability in positional changes, making movement smoother. Smoothposes smooths poses.
[0129] The initial pose includes the initial Z-axis pose (init_zpose), the initial Z-value (init z), and the initial 3D (Init_3D). init_zpose refers to the initial pose or position of the human figure along the Z-axis in 3D space, where the Z-axis represents depth or a direction perpendicular to the screen. init_z refers to the initial Z-axis coordinate value of the human figure in 3D space. Init_3D refers to the initial state of the human figure in 3D space.
[0130] S303. Based on the first portrait trajectory, two-dimensional loss, and smoothing function, optimize the rotation and displacement of the camera parameters to obtain the second portrait trajectory.
[0131] Specifically, in the camera parameter optimization stage, the main focus is on optimizing the camera parameters Rh and Th to optimize the trajectory of the first human figure. This directly affects the rotation and displacement of the root node of the human figure, corresponding to the orientation of the human figure and its position in the world coordinate system. Specifically, optimizing the camera parameters Rh and Th depends on smoothDepth, smoothRh, smoothTh, and 2D_loss.
[0132] S304. Based on the two-dimensional loss, smoothing function and initial pose, optimize the rotation, displacement and pose of the second human figure trajectory to obtain the first three-dimensional motion data.
[0133] Specifically, in the portrait trajectory optimization stage, the definition of the loss term mainly relies on a two-dimensional loss function and smoothing constraints. For optimizing portrait joint rotation, these loss functions aim to smoothly handle sudden joint transitions and optimize the frontal representation of human movement. Specifically, these loss functions can reduce model estimation errors caused by factors such as occlusion, thereby ensuring the continuity and naturalness of joint movements. Furthermore,
[0134] To overcome the depth uncertainty problem caused by over-reliance on 2D keypoints and 2D projections in human pose estimation (especially misjudging forward and backward leaning), an initial pose is introduced. This helps balance the model's dependence on 2D information and enhances the understanding of 3D spatial structure. Specifically, the second human trajectory is optimized, relying on 2D_loss, smoothDepth, smoothRh, smoothTh, Smoothposes, init_zpose, init z, and Init_3D.
[0135] S305. Based on the first data of the three-dimensional motion, obtain the trajectory of the human figure's feet.
[0136] S306. Based on the foot trajectory, the position of the first foot is obtained through a forward dynamics algorithm, and based on the ground contact probability and the position of the first foot, a clustering algorithm is used to construct a predicted ground surface.
[0137] The first foot position includes the toe position and the heel position, so the probability of contact with the ground includes the probability of the toe contacting the ground and the probability of the heel contacting the ground, and the foot speed includes the toe speed and the heel speed.
[0138] S307. Based on the probability of toe contact with the ground, the probability of heel contact with the ground, the speed of toe, and the speed of heel, obtain the sliding loss.
[0139] S308. Grid the foot according to the position of the toe and the position of the heel, and obtain the clipping loss based on the estimated ground and the gridded foot.
[0140] S309. Based on the sliding loss, the clipping loss, and the preset posture loss, the weighted result of each weight combination is obtained by weighted calculation through multiple preset weight combinations.
[0141] S310. Optimize the foot trajectory by selecting the weighted result with the smallest value among multiple weighted results.
[0142] Specifically, the IK algorithm can change the position from the foot to the knee, but it cannot predict how to smoothly adjust a sequence. Even with smoothing through filtering or frame interpolation, robustness cannot be guaranteed in different scenarios. Therefore, by ground contact optimization, which combines ground contact probability with ground prediction, sliding loss and clipping loss are used to solve the foot jump and clipping problems in the subsequent IK algorithm.
[0143] Figure 4 This is a schematic diagram illustrating the principle of ground contact optimization provided in an embodiment of this application. Figure 4 As shown, it is similar to S306-S310, and will not be described again in this embodiment.
[0144] The weighted result with the smallest value among multiple weighted results is represented as:
[0145]
[0146] in, This refers to the initial posture. This refers to shape parameters. This refers to the loss due to sliding. This refers to the loss due to mold penetration. This refers to attitude loss.
[0147]
[0148] in, and These refer to the probability of heel striking the ground and the probability of toe striking the ground, respectively. and These refer to heel speed and toe speed, respectively. When Rh and Th estimates are inaccurate, foot slippage is a common occurrence in human portrait trajectories.
[0149]
[0150] in, This refers to the estimated position of the ground. This refers to the vertex position of the foot grid. This refers to the SMPL parameters at time t (the SMPL model can directly calculate the position of the foot mesh using the provided parameters; here, the average of several mesh nodes on the foot is taken). It refers to the decoder, whose main function is to calculate the foot mesh vertices using SMPL parameters. The main penalty is for clipping and floating when the feet have a high probability of touching the ground.
[0151] S311. Based on the foot trajectory optimized by ground contact, the position of the second foot is obtained through a forward dynamics algorithm, and the end jitter of the second foot position is eliminated through a first filtering algorithm to obtain the end position of the foot.
[0152] S312. Obtain the foot state based on the ground contact probability and the foot end position; wherein, the foot state is used to indicate the ground contact state of the human figure's two feet.
[0153] S313. Based on the foot condition, perform vibration optimization on the foot trajectory after ground contact optimization.
[0154] Specifically Figure 5 This is a schematic diagram illustrating the principle of jitter optimization provided in an embodiment of this application. Figure 5 As shown, the second foot position is first calculated using FK, and the true foot end position is obtained by eliminating end jitter using the first filtering algorithm.
[0155] Secondly, the position of the foot end is classified by the ground contact probability, and the foot state is obtained based on the classification result, expressed as follows.
[0156] Finally, an adjustment strategy is derived based on the foot condition, and the existing Th is adjusted according to the adjustment strategy to achieve vibration optimization.
[0157] In one possible design, the foot position can be one of two feet fully on the ground, two feet completely off the ground, or partially on the ground;
[0158] When the feet are in a position where both feet are fully on the ground, S3053 includes:
[0159] Eliminate foot movement;
[0160] When the feet are in a state where both feet are completely off the ground, S3053 includes:
[0161] By using trajectory acceleration and acceleration constraints, the foot trajectory after ground contact optimization is subjected to jitter optimization;
[0162] When the foot is in partial contact with the ground, S3053 includes:
[0163] The foot trajectory after ground contact optimization is jitter optimized by using a second filtering algorithm and a frame interpolation algorithm.
[0164] Specifically, Table 1 is a state strategy table provided in the embodiments of this application.
[0165] Table 1:
[0166]
[0167] S314. Based on the foot trajectory optimized by shaking, the first three-dimensional motion data is reconstructed using the inverse dynamics algorithm to obtain the second three-dimensional motion data of the human figure.
[0168] Specifically, when reconstructing the first 3D motion data (i.e., the pose parameters in SMPL) based on the jitter-optimized foot trajectory, Th, and Rh, the IK algorithm can be a heuristic algorithm with joint constraints, or other unlisted algorithms; this embodiment does not limit the specific algorithms used. After reconstruction using the IK algorithm, the second 3D motion data of the human figure, i.e., the optimized human figure motion file, can be obtained.
[0169] In other embodiments, the monocular video-based motion capture optimization method can be combined with any human mesh estimation model. It can output optimized, physically accurate human motion files from input monocular video; it can also take existing motion files and video sequences as input and optimize the physical plausibility of the motion separately.
[0170] The technical effects of this application's embodiments also include: by not completely trusting the ground contact detection results, but by using pre-optimization, ground contact optimization, and jitter optimization, the dependence on the ground contact detection results is reduced, and the generalization of 3D motion data reconstruction is improved.
[0171] Figure 6 A flowchart illustrating the model training method provided in the embodiments of this application. Figure 1 .like Figure 6 As shown in the embodiments of this application, the execution entity can be a model training device, which can be located in an electronic device. Figure 1 The second data processing server 130 in the application. The model training method provided in this embodiment includes the following steps:
[0172] S601. Obtain the labeled touch data set.
[0173] The Touch Earth dataset includes multiple training frames, each containing a human image.
[0174] S602. Based on the labeled ground contact dataset, train the ground contact detection model.
[0175] The ground contact detection model provided in this application embodiment is used for Figure 2 and Figure 3An example of a motion capture optimization method based on monocular video. Figure 7 A flowchart illustrating the model training method provided in the embodiments of this application. Figure 2 ,like Figure 7 As shown, the model training method provided in this application embodiment is... Figure 6 Based on the model training method provided in the embodiments, this method is further refined. The model training method provided in this application includes the following steps.
[0176] S701, Obtain the ground contact dataset.
[0177] Specifically, multiple monocular videos (or historical monocular videos) are collected or generated. After initial screening to remove blurry, low-quality, or irrelevant content, a Touch Touch dataset is obtained to ensure its accuracy and effectiveness. The Touch Touch dataset contains high-quality human images with various scenes, angles, and actions to ensure the generalization ability of the Touch Touch detection model. These historical monocular videos can be sourced from publicly available datasets or self-shot videos.
[0178] S702. Based on the ground contact dataset, construct a spring model of the human figure to obtain the force distribution on the feet of the human figure.
[0179] Specifically, a human figure is modeled as a rigid object using a method based on human kinematics. The feet are divided into toes and heels and labeled separately to construct a spring model of the figure. This yields human figure rotation and displacement information that are close to the ground truth, thereby inferring the force on the feet of the figure and obtaining ground contact labels through force label thresholds.
[0180] Furthermore, the acquisition of touch-sensitive tags can be based on the following formula:
[0181]
[0182] in, It refers to the force of contact with the ground (or the ground feedback force). This refers to the Jacobian matrix. This refers to mapping the ground contact force onto the center of mass; It refers to the body's internal force (or joint driving force). This refers to gravitational acceleration. These refer to joint angular velocity and joint acceleration, respectively. These refer to the Coriolis force and centrifugal force vectors, used to represent the additional force effect caused by joint velocity.
[0183] Considering that internal body force is difficult to obtain directly and is generally small in magnitude, we assume it to be zero. Therefore, based on the above formula, the magnitude of the ground contact force, i.e., the force exerted on the foot, can be calculated. Afterward, only a force threshold needs to be applied; therefore, the accuracy of the ground contact force has little impact on the final result.
[0184] S703. Based on the force applied to the feet, obtain the labeled data for each training frame.
[0185] Specifically, based on the foot force calculated in the previous step, each training frame is labeled. The labeled data indicates whether the human figure's foot in the corresponding training frame is in contact with the ground. This is usually achieved by setting a threshold: if the contact force exceeds this threshold, the foot is considered to be in contact with the ground; otherwise, it is considered not to be in contact with the ground. Further, the labeled data indicates the probability of the human figure's foot contacting the ground in the corresponding training frame. The probability of contacting the ground is determined based on the numerical relationship between the contact force and this threshold, such as the ratio of the contact force to the threshold.
[0186] S704. Label the ground contact dataset based on the labeled data of each training frame.
[0187] Specifically, the labeled data is applied to the entire Touch Touch dataset, and each training frame is labeled. In this way, the Touch Touch dataset contains the labeled information for each training frame, namely the probability of the human figure's foot touching the ground.
[0188] S705. Perform multi-view rendering on the labeled ground contact dataset to obtain the key point model.
[0189] Specifically, during video rendering, an engine supporting multi-view rendering can be used. After adding High Dynamic Range (HDR) backgrounds and lighting effects, keypoint estimation is performed. Keypoint estimation includes: for a historical monocular video, multi-view rendering can be performed using cameras with different viewpoints to obtain video sequences from different viewpoints, thereby obtaining keypoint information of the human figure from different viewpoints; keypoint information includes the position, rotation, and trajectory of keypoints. This information is integrated into a keypoint model, which is used to indicate the rotation and trajectory information of the keypoints of the human figure.
[0190] S706. Based on the key point model, intra-frame and inter-frame joint interaction information is obtained through the pre-set DSTformer network and the transformer's attention mechanism.
[0191] Specifically, the ground contact detection model is mainly based on 2D keypoint recognition (vit) and temporal transformer (DSTformer) structure. The 2D keypoints are first projected into high-dimensional spatial features, and then the attention mechanism of the transformer is used to model the joint position relationship at the same time and the joint changes before and after time, respectively capturing the joint interactions within and between frames.
[0192] S707. Based on joint interaction information, a ground contact detection model is trained using a pre-set binary cross-entropy loss function.
[0193] Specifically, the training process employs image pixel-based normalization and uses the binary cross-entropy loss function (BCEloss). BCEloss is a commonly used classification loss function that can be extended to multi-label (multiple touchdown probabilities) classification tasks. Through iterative training, the touchdown detection model learns how to accurately detect human touchdowns from input monocular videos. Ultimately, the trained touchdown detection model can be used for touchdown detection tasks in real-world applications.
[0194] In one possible design, to ensure the probability distribution matches the foot-touching probability and is used in subsequent optimization stages, a small-weighted velocity loss of the foot positions in adjacent frames is added as a supplement to the loss function. Therefore, in each iteration of training, the method also includes:
[0195] S7041. Calculate the first loss function of the ground contact probability output by the ground contact detection model, and the velocity loss function between frames.
[0196] S7042. The second loss function is obtained by weighted calculation based on the first loss function and the velocity loss function.
[0197] The weight of the first loss function is greater than the weight of the velocity loss function.
[0198] S7043. Use the second loss function as the new first loss function.
[0199] The technical effects of this application's embodiments include: by constructing a spring model of a human figure, the labeled data of each training frame of the ground-touch dataset is obtained, and automated labeling is performed, solving the problem of high manual labeling costs for monocular videos; multi-view rendering is performed on the labeled ground-touch dataset to obtain key point information of the human figure from different viewpoints, improving the richness of the dataset; and the ground-touch detection model is trained through the DSTformer network and the transformer's attention mechanism, as well as the binary cross-entropy loss function, improving the accuracy of the ground-touch detection model.
[0200] In other embodiments, the unlabeled touchdown dataset can be rendered from multiple perspectives, and the touchdown detection model can be obtained based on the keypoint model and the labeled data of each training frame.
[0201] Figure 8 This is a schematic diagram illustrating the process of reconstructing the second data of three-dimensional motion according to an embodiment of this application. Figure 8 As shown, the process of reconstructing the second data of three-dimensional motion provided in this application embodiment is... Figures 2 to 5 The embodiment provides a motion capture optimization method based on monocular video, and Figure 6 and Figure 7 Based on the model training method provided in the embodiments, this method is further refined. The process for reconstructing the second data of three-dimensional motion provided in the embodiments of this application includes:
[0202] The first step is to perform keypoint estimation based on the acquired monocular video to obtain a keypoint model. The second step is to obtain a ground contact detection model (or contact-net) based on the keypoint model. The ground contact detection model is used to indicate the probability of the feet of a person in the monocular video touching the ground.
[0203] At the same time, pose estimation is performed based on the monocular video to obtain the first human figure trajectory; then, a pre-optimize module is performed on the first human figure trajectory to obtain the first three-dimensional motion data of the human figure, so as to obtain the foot trajectory of the human figure.
[0204] The second step is to perform foot-optimization on the foot trajectory based on the ground contact probability, including bottom contact optimization and jitter optimization.
[0205] The third step is to reconstruct the first three-dimensional motion data of the human figure based on the optimized foot trajectory using the IK algorithm, thereby obtaining the second three-dimensional motion data of the human figure.
[0206] Figure 9 This is a schematic diagram illustrating the process of labeling the ground contact dataset provided in an embodiment of this application. Figure 9 As shown, the monocular video annotation process provided in this application embodiment is... Figure 6 and Figure 7 Based on the model training method provided in the embodiments, this method is further refined. The monocular video annotation process provided in this application embodiment includes:
[0207] The first step is to construct a spring model of the human figure based on the obtained ground contact dataset to obtain the force distribution on the feet. Then, based on the force distribution, the labeled data for each frame of the monocular video is obtained through a force label threshold. The feet include the toes and heels, so each frame includes four ground contact labels.
[0208] Meanwhile, the Touch Touch dataset is rendered from multiple perspectives (or the Touch Touch dataset after annotation is rendered from multiple perspectives) to obtain a keypoint model; the keypoint model is used to indicate information about multiple keypoints, for example, each training frame includes 17 keypoints.
[0209] The second step is to construct a ground contact detection model based on the labeled data of each training frame and the key point model.
[0210] Figure 10 This is a schematic diagram of the structure of the motion capture optimization device based on monocular video provided in the embodiments of this application, as shown below. Figure 10 As shown in this embodiment, the motion capture optimization device based on monocular video can be located in an electronic device. The motion capture optimization device based on monocular video includes:
[0211] The result determination module 101 is used to acquire monocular video and input the monocular video into a preset ground touch detection model to obtain the ground touch detection result output by the ground touch detection model; wherein, the ground touch detection result is used to indicate the ground touch probability of the human figure's foot in the monocular video; the ground touch detection model is trained by an annotated ground touch dataset, and the annotated ground touch dataset is used to indicate whether the human figure's foot in the training frame touches the ground.
[0212] The trajectory determination module 102 is used to obtain the first three-dimensional motion data of the human figure based on the monocular video, and to obtain the foot trajectory of the human figure based on the first three-dimensional motion data.
[0213] The foot optimization module 103 is used to optimize the foot trajectory based on the ground contact probability, and to optimize the jitter of the optimized foot trajectory based on the ground contact probability.
[0214] The data reconstruction module 104 is used to reconstruct the first three-dimensional motion data of the human figure based on the foot trajectory optimized by shaking, and obtain the second three-dimensional motion data of the human figure.
[0215] The motion capture optimization device based on monocular video provided in this application embodiment can perform... Figure 2 The technical solution of the method embodiment shown has the same implementation principle and technical effect as... Figure 2 The method embodiments shown are similar, and will not be described again in the embodiments of this application.
[0216] Meanwhile, the motion capture optimization device based on monocular video provided in this application embodiment is a further refinement based on the motion capture optimization device based on monocular video provided in the previous application embodiment.
[0217] In one possible design, the foot optimization module 103 includes:
[0218] The ground determination module is used to obtain the position of the first foot through a forward dynamics algorithm based on the foot trajectory, and to construct a predicted ground based on the ground contact probability and the position of the first foot through a clustering algorithm.
[0219] The ground contact optimization module is used to optimize the foot trajectory based on the ground contact probability, the first foot position, the estimated ground, and the foot speed indicated by the foot trajectory.
[0220] In one possible design, the first foot position includes the toe position and the heel position, then the ground contact probability includes the toe contact probability and the heel contact probability, and the foot velocity includes the toe velocity and the heel velocity.
[0221] The ground contact optimization module includes:
[0222] The sliding loss module is used to obtain the sliding loss based on the probability of toe contact with the ground, the probability of heel contact with the ground, the toe speed, and the heel speed.
[0223] The clipping loss module is used to rasterize the foot based on the toe position and heel position, and to obtain the clipping loss based on the estimated ground and the rasterized foot.
[0224] The first weighting module is used to calculate the weighted result of each weight combination based on the sliding loss, the clipping loss and the preset attitude loss through a variety of preset weight combinations.
[0225] The first optimization module is used to optimize the foot trajectory for ground contact based on the weighted result with the smallest value among multiple weighted results.
[0226] In one possible design, the foot optimization module 103 also includes:
[0227] The jitter filtering module is used to obtain the second foot position through a forward dynamics algorithm based on the foot trajectory optimized by ground contact, and to eliminate the jitter at the end of the second foot position through a first filtering algorithm to obtain the foot end position;
[0228] The position determination module is used to obtain the foot state based on the ground contact probability and the foot end position; wherein, the foot state is used to indicate the ground contact state of the human figure's two feet;
[0229] The second optimization module is used to optimize the shaking of the foot trajectory after the ground contact optimization, based on the foot condition.
[0230] In one possible design, the foot position can be one of two feet fully on the ground, two feet completely off the ground, or partially on the ground;
[0231] When both feet are fully on the ground, the second optimization module is used to eliminate foot movement.
[0232] When both feet are off the ground, the second optimization module is used to optimize the shaking of the foot trajectory after ground contact optimization by using trajectory acceleration and acceleration constraints.
[0233] When the foot is partially in contact with the ground, the second optimization module is used to optimize the jitter of the foot trajectory after ground contact optimization by using the second filtering algorithm and the frame interpolation algorithm.
[0234] In one possible design, the trajectory determination module 102 includes:
[0235] The first trajectory module is used to obtain the first human figure trajectory, as well as the human figure's two-dimensional loss, smoothing function, and initial pose, based on the monocular video through pose estimation.
[0236] The second trajectory module is used to optimize the rotation and displacement of camera parameters based on the first portrait trajectory, two-dimensional loss and smoothing function to obtain the second portrait trajectory.
[0237] The third optimization module is used to optimize the rotation, displacement, and posture of the second human figure trajectory based on the two-dimensional loss, smoothing function, and initial posture to obtain the first three-dimensional motion data.
[0238] The motion capture optimization device based on monocular video provided in this application embodiment can perform... Figures 2 to 5 The technical solution of the method embodiment shown has the same implementation principle and technical effect as... Figures 2 to 5 The method embodiments shown are similar, and will not be described again in the embodiments of this application.
[0239] Figure 11 This is a schematic diagram of the structure of the model training device provided in the embodiments of this application, as shown below. Figure 11 As shown in this embodiment, the model training device can be located in an electronic device. The model training device includes:
[0240] The data acquisition module 111 is used to acquire the labeled Touch Touch dataset; wherein, the Touch Touch dataset includes multiple training frames, and each training frame includes a human image;
[0241] The model training module 112 is used to train a ground contact detection model based on the labeled ground contact dataset; wherein the ground contact detection model is used in the motion capture optimization device based on monocular video as described in the above embodiment.
[0242] The model training device provided in this application embodiment can execute... Figure 6 The technical solution of the method embodiment shown has the same implementation principle and technical effect as... Figure 6 The method embodiments shown are similar, and will not be described again in the embodiments of this application.
[0243] Meanwhile, the model training device provided in this application embodiment is further refined based on the model training device provided in the previous application embodiment.
[0244] In one possible design, the model training module 112 includes:
[0245] The view rendering module is used to perform multi-view rendering on the labeled ground touch dataset to obtain the key point model; the key point model is used to indicate the rotation and trajectory information of the key points of the portrait.
[0246] The interaction determination module is used to obtain intra-frame and inter-frame joint interaction information based on the key point model, through the pre-set DSTformer network and the transformer's attention mechanism.
[0247] The model determination module is used to train a ground contact detection model based on joint interaction information using a pre-defined binary cross-entropy loss function.
[0248] In one possible design, the model training device also includes:
[0249] The velocity loss module is used to calculate the first loss function of the ground contact probability output by the ground contact detection model, as well as the velocity loss function between frames;
[0250] The second weighting module is used to calculate the second loss function by weighting the first loss function and the velocity loss function; wherein the weight of the first loss function is greater than the weight of the velocity loss function.
[0251] The function replacement module is used to replace the second loss function with the new first loss function.
[0252] In one possible design, the data acquisition module 111 includes:
[0253] The data acquisition module is used to acquire ground contact datasets;
[0254] The spring construction module is used to build a spring model of a human figure based on the ground contact dataset, so as to obtain the force on the feet of the human figure;
[0255] The annotation determination module is used to obtain annotation data for each training frame based on the force applied to the feet; the annotation data is used to indicate whether the feet of the human figure in the corresponding training frame are touching the ground.
[0256] The data annotation module is used to annotate the touch-down dataset based on the annotation data of each training frame.
[0257] The model training apparatus provided in this application embodiment can execute... Figures 6 to 9 The technical solution of the method embodiment shown has the same implementation principle and technical effect as... Figures 6 to 9 The methods shown in the embodiments are similar, and will not be described again in the embodiments of this application.
[0258] This application also provides an electronic device. Figure 12 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 12 As shown, the electronic device includes at least one processor 121 and a memory 122. The electronic device also includes a communication component 123. The processor 121, memory 122, and communication component 123 are connected via a bus 124.
[0259] In the specific implementation process, at least one processor 121 executes computer execution instructions stored in memory 122, so that at least one processor 121 is used to implement the motion capture optimization method or model training method based on monocular video in the above embodiments.
[0260] The specific implementation process of processor 121 can be found in the above method embodiments, and its implementation principle and technical effect are similar. The embodiments of this application will not be repeated here.
[0261] In the above embodiments, it should be understood that the processor 121 can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.
[0262] The memory 122 may include high-speed RAM memory, and may also include non-volatile memory (NVM), such as at least one disk storage.
[0263] Bus 124 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Bus 124 can be divided into address bus, data bus, control bus, etc. For ease of illustration, the bus 124 in the accompanying drawings of this application is not limited to only one bus or one type of bus.
[0264] The above description addresses the functions implemented by electronic devices and main control devices, and introduces the solutions provided in the embodiments of this application. It is understood that, in order to achieve the above functions, the electronic device or main control device includes hardware structures and / or software modules corresponding to the execution of each function. By combining the units and algorithm steps of the various examples described in the embodiments disclosed in this application, the embodiments of this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the technical solutions of the embodiments of this application.
[0265] This application also provides a computer-readable storage medium storing computer-executable instructions. When executed by a processor, these instructions are used to implement the monocular video-based motion capture optimization method or model training method described above. In the specific implementation of the aforementioned monocular video-based motion capture optimization method or model training method, each module can be implemented as a processor.
[0266] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.
[0267] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in application-specific integrated circuits (ASICs). Alternatively, the processor and the readable storage medium can exist as discrete components in an electronic device or a host device.
[0268] This application also provides a computer program product, including a computer program, which, when executed by a processor, is used to implement the motion capture optimization method or model training method based on monocular video described above.
[0269] The computer program is stored in a readable storage medium, and at least one processor can read the computer program from the readable storage medium and execute the computer program to perform the scheme provided in any of the above embodiments.
[0270] Those skilled in the art will understand that all or part of the steps in the above-described embodiments can be implemented using hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps included in the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0271] The technical solutions of this application have been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it is readily understood by those skilled in the art that the scope of protection of this application is obviously not limited to these specific embodiments. The above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. These modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A motion capture optimization method based on monocular video, characterized in that, The method includes: A monocular video is acquired and input into a preset ground-touch detection model to obtain the ground-touch detection result output by the model. The ground-touch detection result indicates the probability of a person's foot touching the ground in the monocular video. The ground-touch detection model is trained using a labeled ground-touch dataset, which indicates whether a person's foot in the training frames touches the ground. The ground-touch dataset is automatically labeled using a spring model. Based on the monocular video, the first three-dimensional motion data of the human figure is obtained, and based on the first three-dimensional motion data, the foot trajectory of the human figure is obtained. Based on the ground contact probability, the foot trajectory is optimized for ground contact, and based on the ground contact probability and the foot end position, the foot state is obtained. Based on the foot state, the ground contact optimized foot trajectory is further optimized for shaking; wherein, the foot state is used to indicate the ground contact state of the human figure's two feet; Based on the foot trajectory optimized by shaking, the first three-dimensional motion data is reconstructed using an inverse dynamics algorithm to obtain the second three-dimensional motion data of the human figure; The step of optimizing the foot trajectory based on the ground contact probability includes: Based on the foot trajectory, the position of the first foot is obtained through a forward dynamics algorithm, and based on the ground contact probability and the position of the first foot, a clustering algorithm is used to construct an estimated ground surface. The foot trajectory is optimized for ground contact based on the ground contact probability, the first foot position, the estimated ground surface, and the foot speed indicated by the foot trajectory.
2. The method according to claim 1, characterized in that, The first foot position includes the toe position and the heel position, then the ground contact probability includes the toe contact probability and the heel contact probability, and the foot speed includes the toe speed and the heel speed; The step of optimizing the foot trajectory based on the ground contact probability, the first foot position, the estimated ground surface, and the foot speed indicated by the foot trajectory includes: The sliding loss is obtained based on the probability of toe contact with the ground, the probability of heel contact with the ground, the toe speed, and the heel speed; Based on the toe position and the heel position, the foot is gridded, and based on the estimated ground and the gridded foot, the clipping loss is obtained; Based on the sliding loss, the clipping loss, and the preset posture loss, a weighted result for each weight combination is obtained by weighted calculation using a variety of preset weight combinations. The foot trajectory is optimized for ground contact based on the smallest weighted result among the multiple weighted results.
3. The method according to claim 1, characterized in that, The method further includes: Based on the foot trajectory optimized by ground contact, the second foot position is obtained through a forward dynamics algorithm, and the end jitter of the second foot position is eliminated through a first filtering algorithm to obtain the foot end position.
4. The method according to claim 3, characterized in that, The foot position is one of the following: both feet fully on the ground, both feet completely off the ground, and both feet partially on the ground; When the foot position is such that both feet are fully on the ground, the step of optimizing the foot trajectory based on the foot position includes: Eliminate foot movement; When the foot position is such that both feet are off the ground, the step of optimizing the foot trajectory based on the foot position by performing vibration adjustment includes: The foot trajectory after ground contact optimization is subjected to jitter optimization by using trajectory acceleration and acceleration constraints; When the foot is in the partially grounded state, the step of optimizing the foot trajectory based on the ground-touching optimized state includes: The foot trajectory after ground contact optimization is jitter optimized using a second filtering algorithm and a frame interpolation algorithm.
5. The method according to claim 1, characterized in that, The step of obtaining the first three-dimensional motion data of the human figure based on the monocular video includes: Based on the monocular video, the trajectory of the first human figure is obtained through pose estimation, as well as the two-dimensional loss, smoothing function, and initial pose of the human figure; Based on the first portrait trajectory, the two-dimensional loss, and the smoothing function, the rotation and displacement of the camera parameters are optimized to obtain the second portrait trajectory. Based on the two-dimensional loss, the smoothing function, and the initial pose, the rotation, displacement, and pose of the second human figure trajectory are optimized to obtain the first three-dimensional motion data.
6. A model training method, characterized in that, The method includes: Obtain the labeled Touch Touch dataset; wherein the Touch Touch dataset includes multiple training frames, and each training frame includes a human image; A ground touch detection model is trained based on the labeled ground touch dataset; wherein the ground touch detection model is used in the motion capture optimization method based on monocular video as described in any one of claims 1 to 5.
7. The method according to claim 6, characterized in that, The step of training a ground contact detection model based on the labeled ground contact dataset includes: The labeled ground-touching dataset is rendered from multiple perspectives to obtain a keypoint model; wherein, the keypoint model is used to indicate the rotation and trajectory information of the key points of the human figure; Based on the key point model, intra-frame and inter-frame joint interaction information is obtained through a pre-set DSTformer network and a transformer attention mechanism. Based on the joint interaction information, the ground contact detection model is trained using a preset binary cross-entropy loss function.
8. The method according to claim 7, characterized in that, In each iteration of training, the method further includes: Calculate the first loss function of the ground contact probability output by the ground contact detection model, and the velocity loss function between frames; A second loss function is obtained by weighting the first loss function and the velocity loss function; wherein the weight of the first loss function is greater than the weight of the velocity loss function. Use the second loss function as the new first loss function.
9. The method according to claim 6, characterized in that, The obtained labeled touchdown dataset includes: Obtain the aforementioned ground contact dataset; Based on the ground contact dataset, a spring model of the human figure is constructed to obtain the force distribution on the feet of the human figure; Based on the force applied to the feet, annotation data is obtained for each training frame; wherein, the annotation data is used to indicate whether the feet of the human figure in the corresponding training frame are touching the ground; The ground contact dataset is labeled based on the labeled data of each training frame.