Motion capture optimization method and model training method based on monocular video
By adopting a two-stage optimization method in the motion capture technology based on monocular video, the problem of mold penetration and jitter in the three-dimensional reconstruction of the portrait foot is solved, and the accuracy of the three-dimensional motion data is improved.
Patent Information
- Application Number
- CN202510067104.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-15
AI Technical Summary
Motion capture technology based on monocular video is prone to pass through and jitter of portrait feet during the three-dimensional reconstruction process, resulting in a decrease in the accuracy of the three-dimensional motion data.
A two-stage optimization method is adopted: in the first stage, the ground contact probability is obtained through the touchdown detection model, and the ground contact optimization is carried out to solve the problem of jumping and mold penetration of the foot trajectory; in the second stage, the end jitter jitter of the foot trajectory is eliminated through the jitter optimization algorithm to improve the stability of the trajectory.
Through the two-stage optimized foot trajectory, more accurate three-dimensional motion data is reconstructed, improving the accuracy and stability of motion capture.
Smart Images

Figure CN119992651A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a motion capture optimization method and a model training method based on monocular video. Background Art
[0002] Motion capture technology is an interdisciplinary technology that integrates image processing, computer vision and human kinematics. Its core is to capture human motion information and provide a data foundation for many fields such as virtual reality, augmented reality, animation production and medical rehabilitation.
[0003] Among the many motion capture technologies, the motion capture technology based on monocular video has gradually become the research focus due to its economy, convenience and applicability. This technology only needs to use a single camera or a single video stream as information input to extract the key points of the portrait, and then reconstruct the three-dimensional motion data of the portrait through the Inverse Kinematics (IK) algorithm.
[0004] However, due to the lack of depth information, the motion capture technology based on monocular video often encounters the problem of reduced spatial positioning accuracy and occlusion of key points; at the same time, changes in environmental factors such as lighting conditions may also interfere with the extraction of key points. These problems may cause the feet of the portrait to appear through the model and jitter during the 3D reconstruction process, which in turn reduces the accuracy of the 3D motion data reconstructed by the inverse dynamics algorithm. Summary of the invention
[0005] The present application provides a motion capture optimization method and a model training method based on monocular video, which are used to solve the problem that the feet of a portrait may be penetrated and jittered during the three-dimensional reconstruction process, resulting in reduced accuracy of the three-dimensional motion data reconstructed by the inverse dynamics algorithm.
[0006] The first aspect of the present application provides a method for optimizing motion capture based on monocular video, the method comprising:
[0007] Obtain a monocular video, and input the monocular video into a preset ground touchdown detection model to obtain a ground touchdown detection result output by the ground touchdown detection model; wherein the ground touchdown detection result is used to indicate the ground touchdown probability of the feet of the person in the monocular video; the ground touchdown detection model is obtained by training with a labeled ground touchdown data set, and the labeled ground touchdown data set is used to indicate whether the feet of the person in the training frame touch the ground;
[0008] Obtaining first three-dimensional motion data of the portrait according to the monocular video, and obtaining a foot trajectory of the portrait according to the first three-dimensional motion data;
[0009] According to the probability of touching the ground, the foot trajectory is optimized for touching the ground, and according to the probability of touching the ground, the jitter of the foot trajectory after the optimization of touching the ground is optimized;
[0010] According to the foot trajectory after jitter optimization, the three-dimensional motion first data is reconstructed through the inverse dynamics algorithm to obtain the three-dimensional motion second data of the portrait.
[0011] In one possible design, the foot trajectory is optimized based on the ground contact probability, including:
[0012] According to the foot trajectory, the first foot position is obtained through the forward dynamics algorithm, and the estimated ground is constructed through the clustering algorithm based on the ground contact probability and the first foot position;
[0013] The foot trajectory is optimized for touchdown based on the touchdown probability, the first foot position, the estimated ground surface, and the foot speed indicated by the foot trajectory.
[0014] In a possible design, the first foot position includes a toe position and a heel position, the ground contact probability includes a toe contact probability and a heel contact probability, and the foot speed includes a toe speed and a heel speed;
[0015] Optimize the foot trajectory for touchdown based on the touchdown probability, the first foot position, the estimated ground surface, and the foot speed indicated by the foot trajectory, including:
[0016] According to the toe contact probability, heel contact probability, toe speed and heel speed, the sliding loss is obtained;
[0017] The foot is rasterized according to the toe position and the heel position, and the penetration loss is obtained according to the estimated ground and the rasterized foot;
[0018] According to the sliding loss, penetration loss and preset posture loss, a weighted result of each weight combination is obtained through weighted calculation using a preset multiple weight combination;
[0019] The foot trajectory is optimized for ground contact according to a weighted result with the smallest value among the multiple weighted results.
[0020] In a possible design, the foot trajectory after the touchdown optimization is subjected to jitter optimization according to the touchdown probability, including:
[0021] According to the optimized foot trajectory after touching the ground, a second foot position is obtained by a forward dynamics algorithm, and a jitter at the end of the second foot position is eliminated by a first filtering algorithm to obtain a foot end position;
[0022] According to the ground contact probability and the foot end position, the foot state is obtained; wherein the foot state is used to indicate the ground contact state of the feet of the portrait;
[0023] According to the foot state, the foot trajectory after the touchdown optimization is jitter optimized.
[0024] In a possible design, the foot state is one of two feet fully touching the ground, two feet fully off the ground, and partial touching the ground;
[0025] When the foot state is that both feet are completely touching the ground, the foot trajectory after the touchdown optimization is jitter optimized according to the foot state, including:
[0026] Eliminate foot movement;
[0027] When both feet are completely off the ground, the foot trajectory after the touchdown optimization is jitter optimized according to the foot state, including:
[0028] Through trajectory acceleration and acceleration constraints, the foot trajectory after touchdown optimization is jitter optimized;
[0029] When the foot state is partially touching the ground, the foot trajectory after the touchdown optimization is jitter optimized according to the foot state, including:
[0030] The jitter of the foot trajectory after touchdown optimization is optimized through the second filtering algorithm and the interpolation algorithm.
[0031] In a possible design, obtaining first 3D motion data of a person's portrait based on a monocular video includes:
[0032] According to the monocular video, the first portrait trajectory, as well as the portrait's two-dimensional loss, smoothing function and initial pose are obtained through pose estimation;
[0033] According to the first portrait trajectory, the two-dimensional loss and the smoothing function, the rotation and displacement of the camera parameters are optimized to obtain the second portrait trajectory;
[0034] According to the two-dimensional loss, smoothing function and initial posture, the rotation, displacement and posture of the second portrait trajectory are optimized to obtain the first three-dimensional action data.
[0035] A second aspect of the present application provides a model training method, the method comprising:
[0036] Acquire a labeled ground touchdown data set; wherein the ground touchdown data set includes a plurality of training frames, and each training frame includes a human portrait;
[0037] A touchdown detection model is trained based on the labeled touchdown dataset; wherein the touchdown detection model is used in any one of the monocular video-based motion capture optimization methods in the first aspect.
[0038] In a possible design, a touchdown detection model is trained based on the labeled touchdown dataset, including:
[0039] Perform multi-view rendering on the labeled touchdown data set to obtain a key point model; wherein the key point model is used to indicate the rotation information and trajectory information of the key points of the portrait;
[0040] According to the key point model, the joint interaction information within and between frames is obtained through the preset DSTformer network and the attention mechanism of the transformer;
[0041] According to the joint interaction information, the touchdown detection model is trained by a preset binary cross entropy loss function.
[0042] In one possible design, in each iteration of training, the method further includes:
[0043] Calculate a first loss function of the touchdown probability output by the touchdown detection model and a speed loss function between frames;
[0044] A second loss function is obtained by weighted calculation according to the first loss function and the speed loss function; wherein the weight of the first loss function is greater than the weight of the speed loss function;
[0045] The second loss function is used as the new first loss function.
[0046] In a possible design, obtaining a labeled touchdown dataset includes:
[0047] Get the touchdown data set;
[0048] According to the ground contact data set, a spring model of the portrait is constructed to obtain the force of the portrait's feet;
[0049] According to the force applied to the foot, the annotation data of each training frame is obtained; wherein the annotation data is used to indicate whether the foot of the person in the corresponding training frame touches the ground;
[0050] The touchdown dataset is labeled based on the labeled data of each training frame.
[0051] A third aspect of the present application provides a motion capture optimization device based on monocular video, the device comprising:
[0052] A result determination module is used to obtain a monocular video and input the monocular video into a preset ground touchdown detection model to obtain a ground touchdown detection result output by the ground touchdown detection model; wherein the ground touchdown detection result is used to indicate the ground touchdown probability of the feet of the person in the monocular video; the ground touchdown detection model is obtained by training with a labeled ground touchdown data set, and the labeled ground touchdown data set is used to indicate whether the feet of the person in the training frame are touching the ground;
[0053] A trajectory determination module, used to obtain first three-dimensional motion data of a portrait according to a monocular video, and obtain a foot trajectory of the portrait according to the first three-dimensional motion data;
[0054] A foot optimization module, used to optimize the foot trajectory according to the touchdown probability, and to perform jitter optimization on the foot trajectory after the touchdown optimization according to the touchdown probability;
[0055] The data reconstruction module is used to reconstruct the first three-dimensional motion data according to the foot trajectory after jitter optimization through the inverse dynamics algorithm to obtain the second three-dimensional motion data of the portrait.
[0056] A fourth aspect of the present application provides a model training device, the device comprising:
[0057] A data acquisition module is used to acquire a labeled ground contact dataset; wherein the ground contact dataset includes a plurality of training frames, each of which includes a human portrait;
[0058] A model training module is used to train a touchdown detection model based on a labeled touchdown data set; wherein the touchdown detection model is used in the monocular video-based motion capture optimization device of the third aspect.
[0059] A fifth aspect of the present application provides an electronic device, comprising: a memory, and a memory communicatively connected to a processor;
[0060] Memory stores computer-executable instructions;
[0061] When the processor executes the computer-executable instructions stored in the memory, it is used to implement any one of the monocular video-based motion capture optimization methods of the first aspect, or any one of the model training methods of the second aspect.
[0062] The sixth aspect of the present application provides a computer-readable storage medium, which stores computer execution instructions. When the computer execution instructions are executed by a processor, they are used to implement any one of the monocular video-based motion capture optimization methods of the first aspect, or any one of the model training methods of the second aspect.
[0063] The seventh aspect of the present application provides a computer program product, including a computer program, which, when executed by a processor, is used to implement any one of the monocular video-based motion capture optimization methods of the first aspect, or any one of the model training methods of the second aspect.
[0064] The present application provides a method for optimizing motion capture based on monocular video and a method for training a model. The method comprises: inputting a monocular video into a touchdown detection model to obtain a touchdown detection result indicating the touchdown probability of a human foot; obtaining first three-dimensional motion data of the human foot according to the monocular video, and then obtaining the human foot trajectory; optimizing the touchdown of the foot trajectory according to the touchdown probability, and performing jitter optimization on the foot trajectory after the touchdown optimization; reconstructing the first three-dimensional motion data through an inverse dynamics algorithm according to the jitter-optimized foot trajectory, and obtaining second three-dimensional motion data of the human foot. The following technical effects are achieved: optimizing the foot trajectory in two stages, the touchdown optimization in the first stage solves the problem of jump and penetration of the foot trajectory, and the jitter optimization in the second stage solves the problem of jitter at the end of the foot trajectory; reconstructing three-dimensional motion data through an inverse dynamics algorithm according to the foot trajectory after the two-stage optimization, thereby improving the accuracy of the three-dimensional motion data; outputting the touchdown detection result of the monocular video through the touchdown detection model obtained by training the labeled touchdown data set, thereby solving the problem of difficulty in determining the boundary frame of the foot leaving the ground. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0066] Figure 1 A schematic diagram of a scenario provided for an embodiment of the present application;
[0067] Figure 2 Schematic diagram of the process of the motion capture optimization method based on monocular video provided in the embodiment of the present application Figure 1 ;
[0068] Figure 3 Schematic diagram of the process of the motion capture optimization method based on monocular video provided in the embodiment of the present application Figure 2 ;
[0069] Figure 4 A schematic diagram of the principle of ground contact optimization provided by an embodiment of the present application;
[0070] Figure 5 A schematic diagram of the principle of jitter optimization provided in an embodiment of the present application;
[0071] Figure 6 Schematic diagram of the model training method provided in the embodiment of the present application Figure 1 ;
[0072] Figure 7Schematic diagram of the model training method provided in the embodiment of the present application Figure 2 ;
[0073] Figure 8 A schematic diagram of a process of reconstructing second data of a three-dimensional action provided in an embodiment of the present application;
[0074] Fig. 9 A schematic diagram of the process of monocular video annotation provided in an embodiment of the present application;
[0075] Fig.10 A schematic diagram of the structure of a monocular video-based motion capture optimization device provided in an embodiment of the present application;
[0076] Fig.11 A schematic diagram of the structure of a model training device provided in an embodiment of the present application;
[0077] Fig.12 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
[0078] Reference numerals:
[0079] 110 - monocular camera; 120 - first data processing server; 130 - second data processing server;
[0080] 101-result determination module; 102-trajectory determination module; 103-foot optimization module; 104-data reconstruction module;
[0081] 111-data acquisition module; 112-model training module;
[0082] 121 - processor; 122 - memory; 123 - communication component; 124 - bus. DETAILED DESCRIPTION
[0083] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0084] In the present application, words such as "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects. Those skilled in the art will understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit the difference. It should be noted that in the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplary" or "for example" in the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way. In the present application, "at least one" refers to one or more, and "more" refers to two or more.
[0085] It should be noted that the "at..." in this application can be the instant when a certain situation occurs, or it can be a period of time after a certain situation occurs, and this application does not make specific limitations on this. In addition, the motion capture optimization method and model training method based on monocular video provided in this application are only used as examples, and the motion capture optimization method and model training method based on monocular video can also include more or less content. The user information (including but not limited to user device information and user personal information, etc.) and data (including but not limited to data used for analysis, stored data, and displayed data, etc.) involved in one or more embodiments of this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0086] In order to clearly describe the technical solution of this application, some terms and technologies involved in this application are briefly introduced below:
[0087] Motion capture technology: refers to the technology that can extract three-dimensional (3D) human body movements that are consistent with the original video from the video by combining deep learning models and subsequent optimization methods.
[0088] Skinned Multi-Person Linear Model (SMPL): refers to a three-dimensional mesh model used to represent the posture and shape of the human body. SMPL defines a low-dimensional parameter space, including shape parameters and posture parameters. Shape parameters are used to describe the overall shape of the human body, such as height, weight, and body; posture parameters are used to describe the posture of the human body, that is, the rotation angle of the joints.
[0089] Ground contact detection: This refers to identifying the ground contact of the toes and heels of a person’s feet by training a neural network.
[0090] Inverse Kinematics (IK) algorithm: refers to an algorithm that first determines the position of a child bone, then inversely derives the position of its n-level parent bone on the skeleton chain, thereby determining the entire skeleton chain. In this application, the IK algorithm is mainly used for adjusting the posture of the foot.
[0091] Dual-stream Spatio-temporal Transformer (DSTformer) network: refers to a 2D-lifting network with a temporal transformer structure that can capture the temporal characteristics of human motion.
[0092] 2D-lifting network: refers to a 3D key point estimation network that learns prior knowledge about the human body by training the mapping from 2D key points to 3D key points.
[0093] The technical solution of the present application is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The present application will be described below in conjunction with the accompanying drawings.
[0094] In order to clearly understand the technical solution of the present application, the solution of the prior art is first introduced in detail.
[0095] Motion capture technology is an interdisciplinary technology that integrates image processing, computer vision and human kinematics. Its core is to capture human motion information and provide a data basis for many fields such as virtual reality, augmented reality, animation production and medical rehabilitation. Traditional motion capture technology mainly relies on optical motion capture equipment or inertial motion capture equipment. Optical motion capture equipment usually requires multiple high-speed cameras, annotation points and complex computing equipment. It has high accuracy but high cost and requires special venues and environmental conditions. Inertial motion capture equipment relies on multiple sensors attached to the human body to restore human movements by measuring acceleration and angular velocity. It is easily affected by sensor errors and magnetic field interference.
[0096] Among the many motion capture technologies, the motion capture technology based on monocular video has gradually become the research focus due to its economy, convenience and applicability. This technology only needs to use a single camera or a single video stream as information input to extract key points such as joints and skeletons of the portrait, and then reconstruct the three-dimensional motion data of the portrait through the Inverse Kinematics (IK) algorithm.
[0097] However, due to the lack of depth information, motion capture technology based on monocular video often encounters problems such as decreased spatial positioning accuracy and occlusion of key points; at the same time, changes in environmental factors such as lighting conditions may also interfere with the extraction of key points.
[0098] Therefore, in the prior art, the human feet may experience model penetration and jitter during the 3D reconstruction process, which in turn reduces the accuracy of the 3D motion data reconstructed by the inverse dynamics algorithm. In order to solve this problem, the research found that the existing solutions can be divided into two categories. The first type of solution is based on the touchdown detection model to obtain the touchdown result, and the IK algorithm is used to fix the foot position according to the touchdown detection result; the second type of solution is based on the foot speed and other incidental results output by the model as supervision information to optimize the foot trajectory.
[0099] The first problem with the first type of solution is that this type of solution lacks publicly annotated data sets, and it is difficult to determine the boundary frames where the feet leave the ground when annotating the data. Specifically, the data is manually annotated, and it takes about 30 minutes to annotate a short video of about 200 frames. In actual annotation, there is a problem of excessively high manual annotation costs; further, since the distinction between the front and back feet of the portrait in the image is not obvious, especially in some scenes where the feet are not obviously off the ground, it is difficult to determine whether or not the feet have touched the ground. This ambiguity will further increase the annotation cost. Based on this, the present application intends to automatically annotate the ground touchdown data set through a spring model, obtain the ground touchdown label by establishing a spring model for the light capture human body data, and increase the richness of the data set by performing multi-perspective rendering of the light capture human body data through an animation engine.
[0100] The second problem with the first type of solution is that this type of solution involves sudden changes in foot posture that are difficult to optimize when applying the IK algorithm. Specifically, the prior art obtains the boundary frame of the foot leaving the ground, i.e., the touchdown label, by means of a threshold method for the foot speed and height. The model only learns the mapping from key point coordinates to touchdown labels. Once it involves an action type that is not covered by the data set, it cannot be mapped, resulting in the unreliable quality of the labels obtained by this solution and poor generalization of touchdown optimization for various videos. Based on this, the present application intends to use the 2D-lifting network to perform fine-tuning on the touchdown task to ensure a high accuracy rate in various wild scenarios.
[0101] The problem with the second type of solution is that most of these solutions optimize the foot trajectory by designing methods such as foot speed and height loss. However, due to the complexity of the action, the method of optimizing soft constraints alone may lead to insufficient generalization and strong dependence on the touchdown detection results. Errors in a few touchdown detection results will affect the smoothness and naturalness of the entire trajectory. Based on this, the present application intends to optimize the foot trajectory through a two-stage touchdown optimization method that combines optimization loss and IK. The first stage tends to repair the height of the foot trajectory and the height of the foot, and the second stage tends to smooth and fix the foot position.
[0102] Based on the above creative discoveries, the technical solution of the present application is proposed. To solve the above problems, the present application proposes a motion capture optimization method and a model training method based on monocular video. The present application designs a two-stage optimization process, which can optimize and obtain a smooth and physically reasonable three-dimensional motion data of a portrait only through the RGB video sequence of the monocular video and the touchdown detection results output by the touchdown detection model. In addition, the present application also proposes an automatic annotation method for a touchdown dataset, which can generate labels and multi-view video sequences based on the touchdown dataset to increase the richness of the dataset.
[0103] The following introduces the application scenarios of the monocular video-based motion capture optimization method and model training method provided in this application.
[0104] Figure 1 This is a schematic diagram of a scenario provided in an embodiment of the present application. It should be noted that: Figure 1 What is shown are merely examples of scenarios in which the present application can be applied, to help those skilled in the art understand the technical content of the present application, but it does not mean that the present application cannot be used in other devices, systems, environments or scenarios.
[0105] like Figure 1 As shown, the application scenario includes: a monocular camera 110 , a first data processing server 120 and a second data processing server 130 .
[0106] The monocular camera 110 is used to capture the motion video of the motion capture actor and output the captured RGB video, i.e., the monocular video. The monocular camera 110 has high resolution and frame rate, and can clearly record every subtle movement of the motion capture actor.
[0107] The first data processing server 120 is equipped with powerful computing power and sufficient storage space, and is capable of processing and analyzing the monocular video from the monocular camera 110. The first data processing server 120 is in communication connection with the monocular camera 110, and is used to execute the motion capture optimization method based on the monocular video to obtain the second data of the three-dimensional motion of the portrait.
[0108] The second data processing server 130 is also equipped with powerful computing power and sufficient storage space. The second data processing server 130 is used to execute the model training method, can process large-scale data annotation, and train a high-precision touchdown detection model.
[0109] The first data processing server 120 and the second data processing server 130 may be different servers, which are physically separated and each has an independent processor and storage device. The first data processing server 120 and the second data processing server 130 may be located in the same data center or distributed in different geographical locations. After the second data processing server 130 completes the training of the touchdown detection model, it can be transferred to the first data processing server 120 through a file transfer protocol such as FTP (File Transfer Protocol) or SFTP (Secure File Transfer Protocol), or a cloud storage service; the second data processing server 130 can also be deployed as a server that provides a Representational State Transfer API (RESTful API) or a Remote Procedure Call (RPC) service. The first data processing server 120 sends a video frame or feature data to the second data processing server 130 by sending a Hypertext Transfer Protocol (HTTP) request or an RPC call, and receives a response to the touchdown detection result.
[0110] The first data processing server 120 and the second data processing server 130 may also be different logical partitions, containers or virtual machines on the same physical server. The two share the same physical hardware resources, but are logically isolated through virtualization technology. After the second data processing server 130 completes training of the touchdown detection model, the model can be directly stored in a shared storage area such as a network file system (NFS) mount point or a storage area network (SAN) / network attached storage (NAS) storage on the server, and accessed by both as needed; or, the first data processing server 120 and the second data processing server 130 directly exchange touchdown detection results through an inter-process communication mechanism such as shared memory, pipes or sockets.
[0111] Figure 2 Schematic diagram of the process of the motion capture optimization method based on monocular video provided in the embodiment of the present application Figure 1 .like Figure 2As shown, in the embodiment of the present application, the execution subject may be a motion capture optimization device based on monocular video, which may be located in an electronic device, which may be Figure 1 The first data processing server 120 in the embodiment of the present application comprises the following steps:
[0112] S201. Obtain a monocular video, and input the monocular video into a preset ground touchdown detection model to obtain a ground touchdown detection result output by the ground touchdown detection model.
[0113] Specifically, the touchdown detection model is trained with a labeled touchdown dataset, and the labeled touchdown dataset is used to indicate whether the feet of the portrait in the training frame are touching the ground. The touchdown detection model can learn the labeling method of the touchdown probability of the portrait feet by training with the labeled touchdown dataset. Afterwards, the monocular video is input into the touchdown detection model, and the touchdown detection model labels the touchdown probability of the portrait feet in each video frame in the monocular video, and outputs the touchdown detection result. The touchdown detection result is used to indicate the touchdown probability of the portrait feet in the monocular video, specifically, to indicate the touchdown probability of the portrait feet in each video frame in the monocular video.
[0114] S202: Obtain first three-dimensional motion data of the portrait according to the monocular video, and obtain a foot trajectory of the portrait according to the first three-dimensional motion data.
[0115] Specifically, while obtaining the ground contact detection result, the first data processing server also extracts the first three-dimensional motion data of the portrait from the monocular video (the first three-dimensional motion data refers to a type of SMPL data), and then calculates the foot trajectory of the portrait. Among them, the first three-dimensional motion data refers to a data set that describes the motion state of the portrait in three-dimensional space, usually including the three-dimensional coordinates of the portrait joints, motion speed and acceleration, etc., which can fully and accurately reflect the motion characteristics of the portrait.
[0116] It should be noted that the execution order between S201 and S202 is not limited. Figure 2 The execution of S201 first and then S202 is shown, but other execution orders may also be used, such as executing S202 first and then S201, or executing S201 and S202 simultaneously.
[0117] S203: Optimize the foot trajectory according to the touchdown probability, and optimize the jitter of the foot trajectory after the touchdown optimization according to the touchdown probability.
[0118] Specifically, the first-stage touchdown optimization is to solve the jump and penetration problems in the subsequent IK algorithm. Specifically, the IK algorithm can change the position from the foot to the knee, but it cannot predict how to smoothly adjust a sequence. Even if it is smoothed by filtering or interpolation, it still cannot guarantee robustness in different scenarios. Therefore, according to the touchdown probability, the foot trajectory is optimized to solve the jump and penetration problems of the foot trajectory.
[0119] The second stage of jitter optimization first needs to calculate the foot position through the forward kinematics (FK) algorithm, and then eliminate the end jitter of the foot position through filtering. Among them, FK is a physics-based simulation method that can predict the foot trajectory of the portrait based on the initial state of the portrait's position, speed and acceleration, as well as the force state such as gravity and muscle force, combined with the portrait's joint angle and bone length information.
[0120] S204 , reconstructing the first three-dimensional motion data through an inverse dynamics algorithm according to the foot trajectory after the jitter optimization, and obtaining the second three-dimensional motion data of the portrait.
[0121] Specifically, according to the foot trajectory after jitter optimization, the root node trajectory can be adjusted through the IK algorithm, and then the three-dimensional motion first data can be reconstructed according to the adjusted root node trajectory to obtain the final output three-dimensional motion second data of the portrait. Among them, the root node refers to the node at the tail of the portrait spine, and the root node trajectory directly determines the portrait trajectory. Among them, the portrait trajectory refers to the trajectory of the portrait in two-dimensional or three-dimensional space, and the portrait trajectory can be regarded as a special three-dimensional motion data.
[0122] The embodiment of the present application provides a method for optimizing motion capture based on monocular video, the method comprising: inputting the monocular video into a touchdown detection model to obtain a touchdown detection result for indicating the touchdown probability of the feet of a person; obtaining first three-dimensional motion data of the person according to the monocular video, and then obtaining the foot trajectory of the person; optimizing the touchdown of the foot trajectory according to the touchdown probability, and performing jitter optimization on the foot trajectory after the touchdown optimization; reconstructing the first three-dimensional motion data through an inverse dynamics algorithm according to the foot trajectory after the jitter optimization, and obtaining second three-dimensional motion data of the person. The following technical effects are achieved: optimizing the foot trajectory in two stages, the touchdown optimization in the first stage solves the jump and penetration problems of the foot trajectory, and the jitter optimization in the second stage solves the jitter problem at the end of the foot trajectory; reconstructing three-dimensional motion data through an inverse dynamics algorithm according to the foot trajectory after the two-stage optimization, thereby improving the accuracy of the three-dimensional motion data; outputting the touchdown detection result of the monocular video through the touchdown detection model obtained by training the labeled touchdown data set, thereby solving the problem that the boundary frame of the foot leaving the ground is difficult to determine.
[0123] Figure 3 Schematic diagram of the process of the motion capture optimization method based on monocular video provided in the embodiment of the present application Figure 2 ,like Figure 3 As shown, the motion capture optimization method based on monocular video provided in the embodiment of the present application is Figure 2 Based on the monocular video-based motion capture optimization method provided in the embodiment, the method is further refined. The monocular video-based motion capture optimization method provided in the embodiment of the present application includes the following steps.
[0124] S301: Obtain a monocular video, and input the monocular video into a preset ground touchdown detection model to obtain a ground touchdown detection result output by the ground touchdown detection model.
[0125] S302 , obtaining a first portrait trajectory, a two-dimensional loss, a smoothing function, and an initial pose of the portrait through pose estimation according to the monocular video.
[0126] Specifically, the first data of the three-dimensional action can be obtained through the pre-optimization of the first portrait trajectory. The pre-optimization includes two optimization steps: camera parameter optimization and portrait trajectory optimization. Based on the monocular video, the first portrait trajectory, two-dimensional loss, smoothing function and initial pose of the portrait are obtained.
[0127] Two-dimensional loss (2D_loss) refers to the loss function calculated in two-dimensional space.
[0128] Smooth functions include smooth depth (smoothDepth), smooth rotation (smoothRh), smooth displacement (smoothTh) and smooth poses (Smoothposes); smoothDepth refers to a function that smoothes depth data to reduce noise or irregularities and make it smoother or continuous; smoothRh refers to a function that smoothes rotation (Rh) to reduce mutations or jitters in rotation changes and make movement more natural; smoothTh refers to a function that smoothes displacement (Th) to reduce jumps or instability in position changes and make movement smoother; Smoothposes refers to a function that smoothes poses.
[0129] The initial pose includes the initial Z-axis pose (init_zpose), the initial Z value (init z) and the initial three-dimensional (Init_3D); init_zpose refers to the initial pose or position of the portrait along the Z axis in the three-dimensional space, where the Z axis represents the depth or the direction perpendicular to the screen; init_z refers to the initial Z-axis coordinate value of the portrait in the three-dimensional space; Init_3D refers to the initial state of the portrait in the three-dimensional space.
[0130] S303 , optimizing the rotation and displacement of the camera parameters according to the first portrait trajectory, the two-dimensional loss and the smoothing function to obtain a second portrait trajectory.
[0131] Specifically, in the camera parameter optimization stage, the rotation Rh and displacement Th of the camera parameters are mainly optimized to achieve the optimization of the first portrait trajectory, which directly affects the rotation and displacement of the portrait root node, corresponding to the orientation of the portrait and the position in the world coordinate system. Specifically, the optimization of the camera parameters Rh and Th depends on smoothDepth, smoothRh, smoothTh and 2D_loss.
[0132] S304: Optimize the rotation, displacement and posture of the second portrait trajectory according to the two-dimensional loss, the smoothing function and the initial posture to obtain the first three-dimensional motion data.
[0133] Specifically, in the portrait trajectory optimization stage, the definition of the loss term is mainly based on two-dimensional loss functions and smooth constraints. For the optimization of portrait joint rotation, these loss functions are designed to smoothly handle sudden jumps in joints and optimize the positive performance of human body movements. Specifically, through these loss functions, the model estimation error caused by factors such as occlusion can be reduced, thereby ensuring the continuity and naturalness of joint movement. In addition,
[0134] In order to overcome the depth uncertainty problem caused by over-reliance on 2D key points and 2D projections in portrait pose estimation (especially the misjudgment of people leaning forward and backward), the initial pose is introduced, which helps to balance the model's reliance on 2D information and enhance the understanding of 3D spatial structure. Specifically, the second portrait trajectory is optimized, relying on 2D_loss, smoothDepth, smoothRh, smoothTh, Smoothposes, init_zpose, init z, and Init_3D.
[0135] S305: Obtain a foot trajectory of the portrait according to the first three-dimensional motion data.
[0136] S306 , obtaining a first foot position through a forward dynamics algorithm according to the foot trajectory, and constructing an estimated ground surface through a clustering algorithm according to the ground contact probability and the first foot position.
[0137] The first foot position includes a toe position and a heel position, the ground contact probability includes a toe contact probability and a heel contact probability, and the foot speed includes a toe speed and a heel speed.
[0138] S307, obtaining a sliding loss according to the toe-touching probability, the heel-touching probability, the toe speed, and the heel speed.
[0139] S308. Rasterize the foot according to the toe position and the heel position, and obtain the penetration loss according to the estimated ground and the rasterized foot.
[0140] S309: According to the sliding loss, the penetration loss and the preset posture loss, a weighted calculation is performed through a preset plurality of weight combinations to obtain a weighted result of each weight combination.
[0141] S310 , optimizing the foot trajectory according to the weighted result with the smallest value among the multiple weighted results.
[0142] Specifically, the IK algorithm can change the position from the foot to the knee, but it cannot predict how to smoothly adjust a sequence. Even if it is smoothed by filtering or interpolation, it still cannot guarantee robustness in different scenarios. Therefore, through ground contact optimization, that is, the ground contact probability combined with the estimated ground method, the sliding loss and penetration loss are used to solve the foot jump and penetration problems in the subsequent IK algorithm.
[0143] Figure 4 The schematic diagram of the principle of ground contact optimization provided by the embodiment of the present application. Figure 4 As shown, it is similar to S306-S310 and will not be described in detail in this embodiment.
[0144] The weighted result with the smallest value among multiple weighted results is expressed as:
[0145]
[0146] in, is the initial posture, is the shape parameter, is the sliding loss, refers to the penetration loss, It refers to the posture loss.
[0147]
[0148] in, and They refer to the probability of heel touching the ground and the probability of toe touching the ground, and They refer to the heel velocity and toe velocity, respectively. When Rh and Th are estimated inaccurately, foot sliding is common in portrait trajectories.
[0149]
[0150] in, It is the estimated position of the ground. refers to the vertex position of the foot mesh, refers to the SMPL parameters at time t (the SMPL model can directly calculate the position of the foot grid through the provided parameters, here taking the average of several grid nodes of the sole of the foot), Refers to the decoder, whose main function is to calculate the foot grid fixed points through SMPL parameters. It mainly punishes the penetration and floating of the feet when the probability of touching the ground is high.
[0151] S311. According to the optimized foot trajectory after touching the ground, a second foot position is obtained by a forward dynamics algorithm, and the jitter at the end of the second foot position is eliminated by a first filtering algorithm to obtain a foot end position.
[0152] S312: Obtain a foot state according to the ground contact probability and the foot end position; wherein the foot state is used to indicate the ground contact state of both feet of the portrait.
[0153] S313: According to the foot state, perform jitter optimization on the foot trajectory after the ground contact optimization.
[0154] Specifically, Figure 5 The schematic diagram of the principle of jitter optimization provided in the embodiment of the present application. Figure 5 As shown, the second foot position is first calculated by FK, and the jitter of the end is eliminated by the first filtering algorithm to obtain the real foot end position.
[0155] Secondly, the foot end position is classified by the probability of touching the ground, and the foot state is obtained according to the classification result, which is expressed as follows.
[0156] Finally, an adjustment strategy is obtained according to the foot state, and the existing Th is adjusted according to the adjustment strategy to achieve jitter optimization.
[0157] In a possible design, the foot state is one of two feet fully touching the ground, two feet fully off the ground, and partial touching the ground;
[0158] When the foot state is that both feet are completely touching the ground, S3053 includes:
[0159] Eliminate foot movement;
[0160] When the feet are in the state of both feet being off the ground, S3053 includes:
[0161] Through trajectory acceleration and acceleration constraints, the foot trajectory after touchdown optimization is jitter optimized;
[0162] When the foot is in partial contact with the ground, S3053 includes:
[0163] The jitter of the foot trajectory after touchdown optimization is optimized through the second filtering algorithm and the interpolation algorithm.
[0164] Specifically, Table 1 is a status strategy table provided in an embodiment of the present application.
[0165] Table 1:
[0166]
[0167] S314 , reconstructing the first three-dimensional motion data according to the foot trajectory after the jitter optimization by using an inverse dynamics algorithm to obtain second three-dimensional motion data of the portrait.
[0168] Specifically, when reconstructing the first three-dimensional motion data, i.e., the posture parameters in SMPL, based on the foot trajectory after jitter optimization, as well as Th and Rh, the IK algorithm can be a heuristic algorithm with joint constraints, or other unlisted algorithms, which are not limited in this embodiment. After the reconstruction is completed by the IK algorithm, the second three-dimensional motion data of the portrait, i.e., the optimized portrait motion file, can be obtained.
[0169] In other embodiments, the motion capture optimization method based on monocular video can be combined with any human body mesh estimation model. It can output optimized portrait motion files that conform to physical laws through the input monocular video; it can also input existing motion files and video sequences to optimize the physical rationality of the motion separately.
[0170] The technical effects of the embodiments of the present application also include: not completely trusting the touchdown detection results, but reducing the dependence on the touchdown detection results through pre-optimization, touchdown optimization and jitter optimization, and improving the generalization of three-dimensional motion data reconstruction.
[0171] Figure 6 Schematic diagram of the model training method provided in the embodiment of the present application Figure 1 .like Figure 6 As shown, in the embodiment of the present application, the execution subject may be a model training device, which may be located in an electronic device, which may be Figure 1 The second data processing server 130 in the embodiment of the present application comprises the following steps:
[0172] S601: Obtain a labeled touchdown data set.
[0173] The touchdown dataset includes multiple training frames, each of which includes a human image.
[0174] S602: Train a touchdown detection model based on the labeled touchdown dataset.
[0175] The touchdown detection model provided in the embodiment of the present application is used for Figure 2 and Figure 3The embodiment provides a monocular video-based motion capture optimization method. Figure 7 Schematic diagram of the model training method provided in the embodiment of the present application Figure 2 ,like Figure 7 As shown, the model training method provided in the embodiment of the present application is Figure 6 Based on the model training method provided in the embodiment, the model training method provided in the present embodiment is further refined and includes the following steps.
[0176] S701: Acquire a touchdown data set.
[0177] Specifically, multiple monocular videos (or historical monocular videos) are collected or generated, and after preliminary screening, blurry, low-quality or irrelevant content is removed to obtain a touchdown dataset to ensure the accuracy and effectiveness of the touchdown dataset. The touchdown dataset contains portraits with high light capture quality in a variety of scenes, angles and actions to ensure the generalization ability of the touchdown detection model. Among them, these historical monocular videos can come from public datasets or self-shot videos.
[0178] S702: construct a spring model of the portrait according to the ground contact data set to obtain the force condition of the foot of the portrait.
[0179] Specifically, a method based on human kinematics is used to model the portrait as a rigid object, and the feet are divided into toes and heels and labeled separately to construct a spring model of the portrait. The portrait rotation information and portrait displacement information close to the ground truth are obtained, thereby inferring the force condition of the portrait's feet and obtaining the ground contact label through the force label threshold.
[0180] Furthermore, the acquisition of the touchdown label can be based on the following formula:
[0181]
[0182] in, It refers to the ground contact force (or ground feedback force). is the Jacobian matrix, It refers to mapping the contact force to the center of mass; refers to the internal force of the body (or joint driving force), is the acceleration due to gravity, They refer to joint angular velocity and joint acceleration respectively. Refers to the Coriolis and centrifugal force vectors and is used to represent the additional force effects due to joint velocity.
[0183] Considering that the internal force of the body is difficult to obtain directly and the magnitude of the force is generally small, the internal force of the body is assumed to be 0. Therefore, according to the above formula, the magnitude of the ground contact force, that is, the force on the foot, can be calculated. After that, it only needs to pass the force card threshold, so the accuracy of the ground contact force has little impact on the final result.
[0184] S703: Obtaining the labeled data of each training frame according to the force applied to the foot.
[0185] Specifically, each training frame is annotated according to the foot force calculated in the previous step, and the annotated data is used to indicate whether the foot of the portrait in the corresponding training frame touches the ground. This is usually achieved by setting a threshold: if the ground contact force exceeds this threshold, the foot is considered to have touched the ground, otherwise it is considered not to have touched the ground. Furthermore, the annotated data is used to indicate the ground contact probability of the portrait foot in the corresponding training frame, and the ground contact probability is determined according to the numerical relationship between the ground contact force and the threshold, such as the ratio of the ground contact force to the threshold.
[0186] S704: Label the ground contact data set according to the labeled data of each training frame.
[0187] Specifically, the labeled data is applied to the entire touchdown dataset to label each training frame. In this way, the touchdown dataset contains the labeled information of each training frame, that is, the touchdown probability of the foot of the portrait.
[0188] S705 , performing multi-view rendering on the labeled touchdown data set to obtain a key point model.
[0189] Specifically, during the video rendering process, an engine that supports multi-view rendering can be used. After adding a high dynamic range (HDR) background and lighting effects, key point estimation is performed. Key point estimation includes: for a historical monocular video, multi-view rendering can be performed using cameras with different perspectives to obtain video sequences at different perspectives, and then the key point information of the portrait at different perspectives can be obtained; the key point information includes the position, rotation, and trajectory of the key points. This information is integrated into a key point model, which is used to indicate the rotation information and trajectory information of the key points of the portrait;
[0190] S706. According to the key point model, the joint interaction information within and between frames is obtained through the preset DSTformer network and the attention mechanism of the transformer.
[0191] Specifically, the ground contact detection model is mainly based on 2D key point recognition (vit) and temporal transformer (DSTformer) structure. The 2D key points are first projected into high-dimensional spatial features, and then the attention mechanism of the transformer is used to model the joint position relationship at the same time and the joint changes in the previous and next time, respectively, to capture the joint interactions within and between frames.
[0192] S707: According to the joint interaction information, a ground contact detection model is obtained by training with a preset binary cross entropy loss function.
[0193] Specifically, the training adopts a normalization method based on the length and width of the image pixels, and uses the binary cross entropy loss function (BCEloss) for training. BCEloss is a commonly used classification loss function that can be extended to multi-label (multiple touchdown probabilities) classification tasks. Through continuous iterative training, the touchdown detection model learns how to accurately detect the touchdown of the portrait from the input monocular video. Finally, the trained touchdown detection model can be used for touchdown detection tasks in practical applications.
[0194] In one possible design, in order to ensure that the probability distribution can match the foot contact probability and be used in the subsequent optimization stage, a speed loss of the foot position in two adjacent frames with a small weight is added as a supplement to the loss function. Then, in each iteration of training, the method also includes:
[0195] S7041. Calculate a first loss function of the touchdown probability output by the touchdown detection model and a speed loss function between frames.
[0196] S7042. According to the first loss function and the speed loss function, a second loss function is obtained by weighted calculation.
[0197] The weight of the first loss function is greater than the weight of the speed loss function.
[0198] S7043. Use the second loss function as a new first loss function.
[0199] The technical effects of the embodiments of the present application include: by constructing a spring model of the portrait, the labeled data of each training frame of the touchdown data set is obtained, and automatic labeling is performed, thereby solving the problem of high manual labeling cost of monocular video; multi-perspective rendering is performed on the labeled touchdown data set to obtain key point information of the portrait at different perspectives, thereby improving the richness of the data set; the touchdown detection model is obtained through the DSTformer network and the transformer's attention mechanism, as well as the binary cross entropy loss function training, thereby improving the accuracy of the touchdown detection model.
[0200] In other embodiments, the unlabeled touchdown data set may be rendered from multiple perspectives, and a touchdown detection model may be obtained based on the key point model and the labeled data of each training frame.
[0201] Figure 8 The following is a flow chart of the reconstruction of the second data of the three-dimensional action provided in the embodiment of the present application. Figure 8 As shown, the process of reconstructing the second data of the three-dimensional action provided in the embodiment of the present application is Figures 2 to 5 The motion capture optimization method based on monocular video provided in the embodiment, and Figure 6 and Figure 7 Based on the model training method provided in the embodiment, the process of reconstructing the second data of the three-dimensional action provided in the embodiment of the present application is further refined.
[0202] The first step is to perform key point estimation based on the acquired monocular video to obtain a key point model. Secondly, a touchdown detection model (also called contact-net) is obtained based on the key point model. The touchdown detection model is used to indicate the probability of the feet of the person in the monocular video touching the ground.
[0203] At the same time, pose estimation is performed based on the monocular video to obtain the first portrait trajectory. Secondly, the first portrait trajectory is pre-optimized to obtain the first data of the portrait's three-dimensional motion, so as to obtain the portrait's foot trajectory.
[0204] The second step is to perform foot-optimizer on the foot trajectory according to the probability of touching the ground, including bottoming optimization and jitter optimization.
[0205] The third step is to reconstruct the first three-dimensional motion data through the IK algorithm according to the optimized foot trajectory to obtain the second three-dimensional motion data of the portrait.
[0206] Fig. 9 A schematic diagram of the process of labeling a touchdown data set provided in an embodiment of the present application. Fig. 9 As shown, the process of monocular video annotation provided in the embodiment of the present application is Figure 6 and Figure 7 Based on the model training method provided in the embodiment, the process of monocular video annotation provided in the embodiment of the present application includes:
[0207] The first step is to build a spring model of the portrait based on the contact dataset, obtain the force of the portrait's feet, and obtain the label data of each frame of the monocular video based on the force of the feet through the force label threshold. Among them, the foot includes the toe and the heel, so each frame includes four contact labels.
[0208] At the same time, multi-perspective rendering is performed on the touchdown dataset (or multi-perspective rendering can be performed on the above-mentioned labeled touchdown dataset) to obtain a key point model; wherein the key point model is used to indicate information of multiple key points (keypoints), for example, each training frame includes 17 key points.
[0209] The second step is to build a touchdown detection model based on the labeled data of each training frame and the key point model.
[0210] Fig.10 A schematic diagram of the structure of a monocular video-based motion capture optimization device provided in an embodiment of the present application, such as Fig.10 As shown, in the embodiment of the present application, the motion capture optimization device based on monocular video can be located in an electronic device. The motion capture optimization device based on monocular video includes:
[0211] The result determination module 101 is used to obtain a monocular video and input the monocular video into a preset ground touchdown detection model to obtain a ground touchdown detection result output by the ground touchdown detection model; wherein the ground touchdown detection result is used to indicate the ground touchdown probability of the feet of the person in the monocular video; the ground touchdown detection model is obtained by training with a labeled ground touchdown data set, and the labeled ground touchdown data set is used to indicate whether the feet of the person in the training frame touch the ground;
[0212] The trajectory determination module 102 is used to obtain first three-dimensional motion data of the portrait according to the monocular video, and obtain the foot trajectory of the portrait according to the first three-dimensional motion data;
[0213] A foot optimization module 103, configured to optimize the foot trajectory according to the touchdown probability, and to perform jitter optimization on the foot trajectory after the touchdown optimization according to the touchdown probability;
[0214] The data reconstruction module 104 is used to reconstruct the first three-dimensional motion data according to the foot trajectory after the jitter optimization by using the inverse dynamics algorithm to obtain the second three-dimensional motion data of the portrait.
[0215] The monocular video-based motion capture optimization device provided in the embodiment of the present application can be executed Figure 2 The technical solution of the method embodiment shown in the figure has the same implementation principle and technical effect as Figure 2 The method embodiments shown are similar and will not be described in detail in the embodiments of the present application.
[0216] At the same time, the monocular video-based motion capture optimization device provided in the embodiment of the present application is further refined on the basis of the monocular video-based motion capture optimization device provided in the previous application embodiment.
[0217] In a possible design, the foot optimization module 103 includes:
[0218] A ground determination module, used to obtain a first foot position according to the foot trajectory by using a forward dynamics algorithm, and to construct an estimated ground according to the ground contact probability and the first foot position by using a clustering algorithm;
[0219] A touchdown optimization module is used to optimize the foot trajectory for touchdown based on the touchdown probability, the first foot position, the estimated ground surface, and the foot speed indicated by the foot trajectory.
[0220] In a possible design, the first foot position includes a toe position and a heel position, the ground contact probability includes a toe contact probability and a heel contact probability, and the foot speed includes a toe speed and a heel speed;
[0221] Touchdown Optimization Module, including:
[0222] A sliding loss module, used for obtaining sliding loss according to the toe contact probability, the heel contact probability, the toe speed and the heel speed;
[0223] A penetration loss module is used to rasterize the foot according to the toe position and the heel position, and obtain the penetration loss according to the estimated ground and the rasterized foot;
[0224] A first weighting module is used to obtain a weighted result of each weight combination by weighted calculation based on the sliding loss, the penetration loss and the preset posture loss through a preset plurality of weight combinations;
[0225] The first optimization module is used to optimize the foot trajectory according to the weighted result with the smallest value among multiple weighted results.
[0226] In a possible design, the foot optimization module 103 further includes:
[0227] A jitter filtering module, used for obtaining a second foot position by a forward dynamics algorithm according to the optimized foot trajectory after touching the ground, and eliminating the jitter at the end of the second foot position by a first filtering algorithm to obtain a foot end position;
[0228] A position determination module, used to obtain a foot state according to the ground contact probability and the foot end position; wherein the foot state is used to indicate the ground contact state of the feet of the portrait;
[0229] The second optimization module is used to perform jitter optimization on the foot trajectory after the touchdown optimization according to the foot state.
[0230] In a possible design, the foot state is one of two feet fully touching the ground, two feet fully off the ground, and partial touching the ground;
[0231] When the foot state is that both feet are completely touching the ground, the second optimization module is used to eliminate the movement of the foot;
[0232] When the foot state is that both feet are completely off the ground, the second optimization module is used to perform jitter optimization on the foot trajectory after the touchdown optimization through trajectory acceleration and acceleration constraint;
[0233] When the foot is in partial contact with the ground, the second optimization module is used to perform jitter optimization on the foot trajectory after the contact optimization through a second filtering algorithm and an interpolation algorithm.
[0234] In one possible design, the trajectory determination module 102 includes:
[0235] A first trajectory module, used to obtain a first portrait trajectory, a two-dimensional loss, a smoothing function and an initial pose of the portrait through pose estimation according to a monocular video;
[0236] A second trajectory module, used for optimizing the rotation and displacement of camera parameters according to the first portrait trajectory, the two-dimensional loss and the smoothing function, to obtain a second portrait trajectory;
[0237] The third optimization module is used to optimize the rotation, displacement and posture of the second portrait trajectory according to the two-dimensional loss, the smoothing function and the initial posture to obtain the first three-dimensional action data.
[0238] The monocular video-based motion capture optimization device provided in the embodiment of the present application can perform Figures 2 to 5 The technical solution of the method embodiment shown in the figure has the same implementation principle and technical effect as Figures 2 to 5 The method embodiments shown are similar and will not be described in detail in the embodiments of the present application.
[0239] Fig.11 A schematic diagram of the structure of the model training device provided in the embodiment of the present application, such as Fig.11 As shown, in the embodiment of the present application, the model training device may be located in an electronic device. The model training device includes:
[0240] The data acquisition module 111 is used to acquire a labeled ground contact dataset; wherein the ground contact dataset includes a plurality of training frames, each of which includes a human portrait;
[0241] The model training module 112 is used to train a touchdown detection model based on the labeled touchdown data set; wherein the touchdown detection model is used in the motion capture optimization device based on monocular video as in the above-mentioned embodiment.
[0242] The model training device provided in the embodiment of the present application can be executed Figure 6 The technical solution of the method embodiment shown in the figure has the same implementation principle and technical effect as Figure 6 The method embodiments shown are similar and will not be described in detail in the embodiments of the present application.
[0243] At the same time, the model training device provided in the embodiment of the present application is further refined based on the model training device provided in the embodiment of the previous application.
[0244] In one possible design, the model training module 112 includes:
[0245] A perspective rendering module is used to perform multi-perspective rendering on the annotated touchdown data set to obtain a key point model; wherein the key point model is used to indicate the rotation information and trajectory information of the key points of the portrait;
[0246] The interaction determination module is used to obtain the joint interaction information within and between frames based on the key point model through the preset DSTformer network and the attention mechanism of the transformer;
[0247] The model determination module is used to obtain a ground contact detection model through training using a preset binary cross entropy loss function according to joint interaction information.
[0248] In a possible design, the model training device further includes:
[0249] A speed loss module, used to calculate a first loss function of a touchdown probability output by a touchdown detection model and a speed loss function between frames;
[0250] A second weighting module is used to obtain a second loss function by weighted calculation according to the first loss function and the speed loss function; wherein the weight of the first loss function is greater than the weight of the speed loss function;
[0251] Function replacement module, used to use the second loss function as the new first loss function.
[0252] In a possible design, the data acquisition module 111 includes:
[0253] A data acquisition module, used to obtain a touchdown data set;
[0254] A spring construction module is used to construct a spring model of the portrait based on the ground contact data set to obtain the force conditions of the portrait's feet;
[0255] A labeling determination module is used to obtain labeling data for each training frame according to the force applied to the foot; wherein the labeling data is used to indicate whether the foot of the person in the corresponding training frame touches the ground;
[0256] The data annotation module is used to annotate the touchdown data set according to the annotation data of each training frame.
[0257] The model training device provided in the embodiment of the present application can execute Figures 6 to 9 The technical solution of the method embodiment shown in the figure has the same implementation principle and technical effect as Figures 6 to 9 The method embodiments shown are similar and will not be described in detail in the embodiments of the present application.
[0258] The present application also provides an electronic device, Fig.12 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Fig.12 As shown, the electronic device includes: at least one processor 121 and a memory 122. The electronic device also includes a communication component 123. The processor 121, the memory 122 and the communication component 123 are connected via a bus 124.
[0259] During the specific implementation process, at least one processor 121 executes the computer execution instructions stored in the memory 122, so that at least one processor 121 is used to implement the motion capture optimization method or model training method based on monocular video in the above-mentioned embodiment.
[0260] The specific implementation process of the processor 121 can be found in the above-mentioned method embodiment, and its implementation principle and technical effect are similar, so the embodiments of the present application will not be repeated here.
[0261] In the above embodiments, it should be understood that the processor 121 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. The general-purpose processor may be a microprocessor or the processor may be any conventional processor, etc. The steps of the method disclosed in the application may be directly embodied as being executed by a hardware processor, or may be executed by a combination of hardware and software modules in the processor.
[0262] The memory 122 may include a high-speed RAM memory, and may also include a non-volatile storage NVM, such as at least one disk storage.
[0263] The bus 124 may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus 124 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus 124 in the drawings of the present application is not limited to only one bus or one type of bus.
[0264] The above functions implemented by the electronic device and the main control device introduce the scheme provided by the embodiment of the present application. It is understandable that in order to implement the above functions, the electronic device or the main control device includes a hardware structure and / or software module corresponding to each function. In combination with the units and algorithm steps of each example described in the embodiment disclosed in the embodiment of the present application, the embodiment of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the technical solution of the embodiment of the present application.
[0265] The embodiment of the present application also provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the above-mentioned embodiment of the motion capture optimization method or model training method based on monocular video. In the specific implementation of the above-mentioned motion capture optimization method or model training method based on monocular video, each module can be implemented as a processor.
[0266] The above-mentioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The readable storage medium can be any available medium that can be accessed by a general or special-purpose computer.
[0267] An exemplary readable storage medium is coupled to a processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in an electronic device or a main control device.
[0268] An embodiment of the present application also provides a computer program product, including a computer program, which, when executed by a processor, is used to implement the monocular video-based motion capture optimization method or model training method of the above-mentioned embodiment.
[0269] The computer program is stored in a readable storage medium. At least one processor can read the computer program from the readable storage medium. At least one processor executes the computer program to execute the solution provided in any of the above embodiments.
[0270] A person skilled in the art can understand that all or part of the steps of implementing the above-mentioned application embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, the steps of the above-mentioned method embodiment are executed; and the aforementioned storage medium includes: ROM, RAM, disk or optical disk, etc., various media that can store program codes.
[0271] So far, the technical solution of the present application has been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present application is obviously not limited to these specific embodiments, and the above embodiments are only used to illustrate the technical solution of the present application rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A motion capture optimization method based on monocular video, characterized in that: The method comprises: Acquire a monocular video, and input the monocular video into a preset ground touchdown detection model to obtain a ground touchdown detection result output by the ground touchdown detection model; wherein the ground touchdown detection result is used to indicate the ground touchdown probability of the feet of the person in the monocular video; the ground touchdown detection model is obtained by training with a labeled ground touchdown data set, and the labeled ground touchdown data set is used to indicate whether the feet of the person in the training frame touch the ground; Obtaining first three-dimensional motion data of a person's portrait according to the monocular video, and obtaining a foot trajectory of the person's portrait according to the first three-dimensional motion data; According to the ground contact probability, the foot trajectory is optimized for ground contact, and according to the ground contact probability, the foot trajectory after ground contact optimization is optimized for jitter; According to the foot trajectory after jitter optimization, the three-dimensional motion first data is reconstructed by an inverse dynamics algorithm to obtain the three-dimensional motion second data of the portrait.
2. The method according to claim 1, characterized in that: The step of optimizing the foot trajectory according to the ground contact probability includes: According to the foot trajectory, a first foot position is obtained by a forward dynamics algorithm, and according to the ground contact probability and the first foot position, an estimated ground surface is constructed by a clustering algorithm; The foot trajectory is optimized for a touchdown according to the touchdown probability, the first foot position, the estimated ground surface, and a foot speed indicated by the foot trajectory.
3. The method according to claim 2, characterized in that The first foot position includes a toe position and a heel position, the ground contact probability includes a toe contact probability and a heel contact probability, and the foot speed includes a toe speed and a heel speed; The step of optimizing the foot trajectory according to the ground contact probability, the first foot position, the estimated ground surface, and the foot speed indicated by the foot trajectory includes: Obtaining a sliding loss according to the toe contact probability, the heel contact probability, the toe speed, and the heel speed; rasterizing the foot according to the toe position and the heel position, and obtaining a penetration loss according to the estimated ground and the rasterized foot; According to the sliding loss, the penetration loss and the preset posture loss, a weighted result of each weight combination is obtained by weighted calculation through a preset plurality of weight combinations; The foot trajectory is optimized for ground contact according to a weighted result with a minimum value among the multiple weighted results.
4. The method according to claim 1, characterized in that: The step of performing jitter optimization on the foot trajectory after the touchdown optimization according to the touchdown probability includes: According to the optimized foot trajectory after touching the ground, a second foot position is obtained by a forward dynamics algorithm, and a jitter at the end of the second foot position is eliminated by a first filtering algorithm to obtain a foot end position; According to the ground contact probability and the foot end position, a foot state is obtained; wherein the foot state is used to indicate the ground contact state of the feet of the portrait; According to the foot state, jitter optimization is performed on the foot trajectory after the ground contact optimization.
5. The method according to claim 4, characterized in that The foot state is one of both feet completely touching the ground, both feet completely off the ground, and partially touching the ground; When the foot state is that both feet are completely touching the ground, performing jitter optimization on the foot trajectory after the ground contact optimization according to the foot state includes: Eliminate foot movement; When the foot state is that both feet are completely off the ground, performing jitter optimization on the foot trajectory after the ground contact optimization according to the foot state includes: Performing jitter optimization on the foot trajectory after the touchdown optimization through trajectory acceleration and acceleration constraint; When the foot state is the partial ground contact, performing jitter optimization on the foot trajectory after the ground contact optimization according to the foot state includes: The foot trajectory after the touchdown optimization is jitter-optimized by using a second filtering algorithm and a frame interpolation algorithm.
6. The method according to claim 1, characterized in that The step of obtaining first three-dimensional motion data of a person's portrait according to the monocular video includes: According to the monocular video, a first portrait trajectory, as well as a two-dimensional loss, a smoothing function, and an initial pose of the portrait are obtained through pose estimation; According to the first portrait trajectory, the two-dimensional loss and the smoothing function, the rotation and displacement of the camera parameters are optimized to obtain a second portrait trajectory; The rotation, displacement and posture of the second portrait trajectory are optimized according to the two-dimensional loss, the smoothing function and the initial posture to obtain the three-dimensional motion first data.
7. A model training method, characterized in that: The method comprises: Acquire a labeled ground contact data set; wherein the ground contact data set includes a plurality of training frames, each of the training frames includes a human portrait; A touchdown detection model is trained based on the labeled touchdown data set; wherein the touchdown detection model is used in the monocular video-based motion capture optimization method as described in any one of claims 1 to 6.
8. The method according to claim 7, characterized in that The training of a touchdown detection model according to the labeled touchdown dataset includes: Performing multi-view rendering on the labeled touchdown data set to obtain a key point model; wherein the key point model is used to indicate rotation information and trajectory information of key points of the portrait; According to the key point model, the joint interaction information within and between frames is obtained through the preset DSTformer network and the attention mechanism of the transformer; The touchdown detection model is obtained by training according to the joint interaction information through a preset binary cross entropy loss function.
9. The method according to claim 8, characterized in that In each iteration of training, the method further comprises: Calculating a first loss function of the touchdown probability output by the touchdown detection model and a speed loss function between frames; A second loss function is obtained by weighted calculation according to the first loss function and the speed loss function; wherein the weight of the first loss function is greater than the weight of the speed loss function; The second loss function is used as the new first loss function.
10. The method according to claim 7, characterized in that The step of obtaining the labeled ground contact dataset includes: acquiring the touchdown data set; According to the ground contact data set, a spring model of the portrait is constructed to obtain the force condition of the foot of the portrait; According to the force applied to the foot, the annotation data of each training frame is obtained; wherein the annotation data is used to indicate whether the foot of the person in the corresponding training frame touches the ground; The touchdown data set is labeled according to the labeled data of each training frame.
Citation Information
Patent Citations
Motion capture data processing method and device, equipment and storage medium
CN116051699A
Monocular video-based multi-stage human motion capture method and device, and medium
CN116386141A
Motion capture method and system based on parameterized model
CN117541646A
Method and system for processing character motion trail in monocular video motion capture
CN118486074A
6-dof tracking using visual cues
EP4394561A2
Cited By
Three-dimensional human body action generation method and device based on parallel multi-granularity Transform
CN122176200A
A three-dimensional human body motion generation method and device based on a parallel multi-granularity transformer
CN122176200B