Motion capture data optimization method and device, electronic equipment and storage medium

By iteratively updating body shape and pose parameters, optimizing video motion capture data, solving the problem of unreasonable motion data, achieving efficient action reconstruction and reducing repair costs.

CN120496166APending Publication Date: 2025-08-15DUOYI NETWORK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510468518.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the existing video motion capture technology, the action data lacks physicality and requires manual adjustment, resulting in low action output efficiency and high repair costs.

Method used

By obtaining video frame data, using the parameterized human body model and human body pose prior distribution model, iteratively updates the body shape parameters and pose parameters, so that the latent variable distribution meets the preset conditions and optimizes the three-dimensional joint node data.

Benefits of technology

The quality of motion capture data is improved, the action reconstruction is no longer distorted, the action output efficiency is improved, and the repair cost is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496166A_ABST
    Figure CN120496166A_ABST
Patent Text Reader

Abstract

The invention relates to a motion capture data optimization method and device, electronic equipment and a storage medium, and the method comprises the steps: updating a first body type parameter and a first attitude parameter in iteration; the first latent variable distribution obtained on the basis of the updated first body shape parameter and the updated first posture parameter conforms to the preset second latent variable distribution, so that the human body action reconstructed on the basis of the updated first posture parameter is located in the normal activity range of the human body, and the unreasonability of the motion capture data is avoided; and the quality of motion capture data is improved. As the reconstructed human body action is not distorted any more, manual adjustment is not needed, the action output efficiency is improved, and the repair cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video motion capture, and in particular to a motion capture data optimization method, device, electronic equipment and storage medium. Background Art

[0002] Motion capture is a technology that records and reproduces the movements of objects or people in the real world. It is commonly used in game development, film special effects, sports analysis and other fields. It captures the positional changes of actors or objects and then uses this data to reconstruct and reproduce these movements in a digital environment.

[0003] Motion capture technology mainly includes three different methods: inertial capture, optical capture, and video capture. Among them, video motion capture is a method that uses a video camera to record human movement and extract motion data through computer vision technology.

[0004] However, in the existing technology, the motion data of video motion capture lacks physicality, is not realistic enough, is difficult to use directly, and requires manual adjustment, which seriously restricts the motion output efficiency of video motion capture and increases the restoration cost. Summary of the Invention

[0005] Based on this, the purpose of the present invention is to provide a motion capture data optimization method, device, electronic device and storage medium, which have the advantages of improving the quality of motion capture data, increasing action output efficiency and reducing repair costs.

[0006] According to a first aspect of an embodiment of the present application, a method for optimizing motion capture data is provided, comprising the following steps:

[0007] Obtaining video frame data; wherein the video frame data is used to record human body movements;

[0008] Obtaining first three-dimensional joint point data according to the video frame data;

[0009] Obtaining first body shape parameters and first posture parameters; obtaining second three-dimensional joint data based on the first body shape parameters, the first posture parameters, and the parameterized human body model; obtaining a first latent variable distribution based on the first body shape parameters, the first posture parameters, and the human body posture prior distribution model;

[0010] Iteratively updating the first body shape parameter and the first posture parameter according to the first three-dimensional joint point data, the second three-dimensional joint point data, the first latent variable distribution, and the preset second latent variable distribution until a preset condition is satisfied, thereby obtaining the second body shape parameter and the second posture parameter;

[0011] Optimized three-dimensional joint point data is obtained according to the second body shape parameter, the second posture parameter, and the parameterized human body model.

[0012] According to a second aspect of an embodiment of the present application, a motion capture data optimization device is provided, comprising:

[0013] A video frame data acquisition module is used to acquire video frame data; wherein the video frame data is used to record human body movements;

[0014] A joint point data acquisition module, configured to obtain first three-dimensional joint point data based on the video frame data;

[0015] a latent variable distribution acquisition module for acquiring a first body shape parameter and a first posture parameter; obtaining second three-dimensional joint point data based on the first body shape parameter, the first posture parameter, and a parameterized human body model; and obtaining a first latent variable distribution based on the first body shape parameter, the first posture parameter, and a human body posture prior distribution model;

[0016] a posture parameter acquisition module, configured to iteratively update the first body shape parameter and the first posture parameter according to the first three-dimensional joint point data, the second three-dimensional joint point data, the first latent variable distribution, and the preset second latent variable distribution until a preset condition is satisfied, thereby obtaining the second body shape parameter and the second posture parameter;

[0017] The joint point data optimization module is used to obtain optimized three-dimensional joint point data based on the second body shape parameter, the second posture parameter and the parameterized human body model.

[0018] According to a third aspect of an embodiment of the present application, an electronic device is provided, comprising: a processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing any one of the motion capture data optimization methods described above.

[0019] According to a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the motion capture data optimization method as described above is implemented.

[0020] The embodiment of the present application obtains video frame data; wherein the video frame data is used to record human body movements; obtains first three-dimensional joint point data based on the video frame data; obtains first body shape parameters and first posture parameters; obtains second three-dimensional joint point data based on the first body shape parameters, the first posture parameters and a parameterized human body model; obtains a first latent variable distribution based on the first body shape parameters, the first posture parameters and a human body posture prior distribution model; iteratively updates the first body shape parameters and the first posture parameters based on the first three-dimensional joint point data, the second three-dimensional joint point data, the first latent variable distribution and a preset second latent variable distribution until a preset condition is met, thereby obtaining second body shape parameters and second posture parameters; and obtains optimized three-dimensional joint point data based on the second body shape parameters, the second posture parameters and the parameterized human body model. The present application updates the first body shape parameters and the first posture parameters in the iteration so that the first latent variable distribution obtained based on the updated first body shape parameters and the updated first posture parameters conforms to the preset second latent variable distribution, thereby ensuring that the human body movements reconstructed based on the updated first posture parameters are within the normal range of human activities, avoiding the irrationality of motion capture data and improving the quality of motion capture data. Since the reconstructed human body movements are no longer distorted and do not require manual adjustment, the efficiency of movement output is improved and the cost of restoration is reduced.

[0021] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application.

[0022] For better understanding and implementation, the present invention is described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 A flowchart of a motion capture data optimization method provided in one embodiment of the present application;

[0024] Figure 2 A structural block diagram of a motion capture data optimization device provided in one embodiment of the present application;

[0025] Figure 3 A schematic block diagram of the structure of an electronic device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0026] In order to make the objectives, technical solutions and advantages of the present application clearer, the embodiments of the present application will be described in further detail below with reference to the accompanying drawings.

[0027] It should be clear that the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0028] The terms used in the embodiments of the present application are for the purpose of describing specific embodiments only and are not intended to limit the embodiments of the present application. The singular forms "a," "the," and "the" used in the embodiments of the present application and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.

[0029] When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims. In the description of the present application, it should be understood that the terms "first", "second", "third", etc. are only used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence, nor can they be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to the specific circumstances.

[0030] In addition, in this application, unless otherwise specified, "plurality" refers to two or more. "And / or" describes the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. The character " / " generally indicates that the associated objects are in an "or" relationship.

[0031] The motion capture data optimization method provided in the embodiments of the present application can be performed by a motion capture data optimization device. The motion capture data optimization device can be implemented via software and / or hardware. The motion capture data optimization device can be composed of two or more physical entities, or a single physical entity. The motion capture data optimization device can be any electronic device equipped with image processing software, such as a computer, mobile phone, tablet, or other smart device.

[0032] During the development of this invention, the inventors discovered that the human motion reconstructed from video motion capture data in related technologies exceeds the normal range of human movement. When applied to game models, these models appear distorted and require manual adjustment. Manual adjustment is time-consuming, inefficient, requires specialized personnel, and is costly.

[0033] To this end, the present application updates the first body shape parameters and the first posture parameters during iterations, so that the first latent variable distribution obtained based on the updated first body shape parameters and the updated first posture parameters conforms to the preset second latent variable distribution. This ensures that the human motion reconstructed based on the updated first posture parameters falls within the normal range of human activity, avoiding irrationality in the motion capture data and improving its quality. Because the reconstructed human motion is no longer distorted, manual adjustment is not required, reducing restoration costs.

[0034] See also Figure 1 , which is a flow chart of a motion capture data optimization method provided in one embodiment of the present application. The motion capture data optimization method provided in the embodiment of the present application comprises the following steps:

[0035] S10: Obtain video frame data; wherein the video frame data is used to record human body movements.

[0036] In the embodiment of the present application, a human body action is captured by a camera to obtain an action video, and all video frames of the action video are used as video frame data.

[0037] Optionally, multiple cameras at different positions are used to capture human body movements to obtain multiple action videos from different perspectives, wherein one camera corresponds to one action video, each action video includes several video frames, and all video frames of all action videos are used as video frame data.

[0038] S20: Obtain first three-dimensional joint point data according to the video frame data.

[0039] The first three-dimensional joint point data includes the three-dimensional spatial position coordinates of each joint point of the human body.

[0040] In an embodiment of the present application, each joint point in a video frame is identified to obtain the two-dimensional position coordinates of each joint point. The two-dimensional position coordinates of each joint point are converted into three-dimensional space coordinates using the intrinsic and extrinsic parameters of the camera to obtain the first three-dimensional joint point data. Among them, the intrinsic parameters of the camera are used to describe the properties of the camera itself, including focal length, principal point position, pixel size, and distortion coefficient. The extrinsic parameters of the camera are used to describe the position and orientation of the camera in the world coordinate system, that is, the spatial relationship of the camera relative to the scene.

[0041] S30: Obtain first body parameters and first posture parameters; obtain second three-dimensional joint data based on the first body parameters, the first posture parameters and the parameterized human body model; obtain a first latent variable distribution based on the first body parameters, the first posture parameters and the human body posture prior distribution model.

[0042] A parametric human body model is one that uses body shape and posture parameters to define the shape and posture of the human body. By adjusting the body shape and posture parameters, human body models of different body shapes and postures can be generated.

[0043] The first body shape parameter and the first posture parameter are randomly initialized body shape parameters and posture parameters. The body shape parameters are used to describe the shape and size of the human body. The posture parameters are used to describe the spatial position and orientation of the human body, including the joint rotation angles of each joint point.

[0044] Among them, the human posture prior distribution model is a deep learning model with a VAE structure. VAE (Variational Autoencoder) is a generative model that aims to learn the latent space distribution of input data. VAE includes an encoder and a decoder. The encoder maps the input data (for example, body shape parameters and posture parameters) to the latent space, learns the distribution of the latent variables, and outputs the mean and variance of the latent variables to form a set of high-dimensional latent vectors. The decoder reconstructs the input data from the samples in the latent space, for example, converting the latent variables into the original human body shape and posture parameters. The human posture prior distribution model provides a constraint mechanism to produce reasonable and accurate human posture predictions in the presence of noise, occlusion or incomplete data.

[0045] In an embodiment of the present application, the body shape parameters and posture parameters of a parameterized human body model are initialized to obtain first body shape parameters and first posture parameters. Based on the first body shape parameters and first posture parameters, the parameterized human body model is adjusted to obtain the three-dimensional position coordinates of each joint. The first body shape parameters and first posture parameters are input into a human body posture prior distribution model. The encoder of the human body posture prior distribution model outputs the mean and variance of the latent variable to obtain a first latent variable distribution.

[0046] S40: Iteratively update the first body shape parameter and the first posture parameter according to the first three-dimensional joint point data, the second three-dimensional joint point data, the first latent variable distribution, and the preset second latent variable distribution until the preset conditions are met to obtain the second body shape parameter and the second posture parameter.

[0047] The preset second latent variable distribution is set according to actual needs and can be a Gaussian distribution.

[0048] The preset condition may be that the number of iterative updates reaches a preset number.

[0049] In an embodiment of the present application, the position error between the second three-dimensional joint point data and the first three-dimensional joint point data, the mean error and variance error between the first latent variable distribution and the preset second latent variable distribution are calculated, and based on the position error, mean error and variance error, the gradient descent algorithm is used to update the first body shape parameter and the first posture parameter until the preset conditions are met, thereby obtaining the second body shape parameter and the second posture parameter.

[0050] S50: Obtain optimized three-dimensional joint point data according to the second body shape parameter, the second posture parameter, and the parameterized human body model.

[0051] In an embodiment of the present application, the parameterized human body model is adjusted according to the second body shape parameter and the second posture parameter to obtain the optimized three-dimensional position coordinates of each joint point.

[0052] The embodiment of the present application is applied, by obtaining video frame data; wherein the video frame data is used to record human body movements; obtaining first three-dimensional joint point data based on the video frame data; obtaining first body shape parameters and first posture parameters; obtaining second three-dimensional joint point data based on the first body shape parameters, the first posture parameters and a parameterized human body model; obtaining a first latent variable distribution based on the first body shape parameters, the first posture parameters and a human body posture prior distribution model; iteratively updating the first body shape parameters and the first posture parameters based on the first three-dimensional joint point data, the second three-dimensional joint point data, the first latent variable distribution and a preset second latent variable distribution until a preset condition is met, thereby obtaining second body shape parameters and second posture parameters; obtaining optimized three-dimensional joint point data based on the second body shape parameters, the second posture parameters and the parameterized human body model. The present application updates the first body shape parameters and the first posture parameters in the iteration so that the first latent variable distribution obtained based on the updated first body shape parameters and the updated first posture parameters conforms to the preset second latent variable distribution, thereby ensuring that the human body movements reconstructed based on the updated first posture parameters are within the normal range of human activities, avoiding the irrationality of motion capture data and improving the quality of motion capture data. Since the reconstructed human body movements are no longer distorted and do not require manual adjustment, the efficiency of movement output is improved and the cost of restoration is reduced.

[0053] In one embodiment, the video frame data includes a plurality of video frames of human body movements captured by multiple cameras. Step S20 includes steps S21 to S23, which are specifically as follows:

[0054] S21: Extract the two-dimensional position information of each joint point from each video frame.

[0055] In an embodiment of the present application, each video frame is input into a trained deep learning model to obtain the two-dimensional position coordinates of each joint point in each video frame. The trained deep learning model is trained based on the video frame data and the annotated joint point data.

[0056] S22: Obtain three-dimensional position information of each joint point based on the two-dimensional position information and parameter information of the corresponding camera.

[0057] The camera parameter information includes the calibrated camera intrinsic parameters and camera extrinsic parameters.

[0058] In this embodiment of the present application, for the kth joint point in the jth video frame captured by the i-th camera, the two-dimensional position coordinates of the kth joint point are obtained. Using the intrinsic and extrinsic parameters of the i-th camera, the two-dimensional position coordinates of the kth joint point in the image coordinate system are converted to three-dimensional position coordinates in the world coordinate system. This operation is repeated for each joint point to obtain the three-dimensional position coordinates of each joint point.

[0059] S23: Performing stereo triangulation processing on the three-dimensional position information of each joint point to obtain first three-dimensional joint point data of each joint point.

[0060] Among them, stereo triangulation refers to a method of using two or more cameras to capture images of the same scene from different perspectives, and then calculating the three-dimensional position of objects in the scene by matching corresponding points in the images.

[0061] In an embodiment of the present application, taking the kth joint point in the jth video frame as an example, the three-dimensional position coordinates corresponding to the kth joint point in the jth video frame taken by each camera are obtained, these three-dimensional position coordinates are stereo triangulated to obtain a three-dimensional position coordinate, and the three-dimensional position coordinate is used as the three-dimensional position coordinate of the kth joint point in the jth video frame.

[0062] The embodiment of the present application utilizes video frame data captured by multiple cameras and combines it with stereo triangulation to improve the accuracy of the three-dimensional position coordinates of each joint point.

[0063] In one embodiment, the step of obtaining the second three-dimensional joint point data according to the first body shape parameter, the first posture parameter, and the parameterized human body model in step S30 includes steps S31 to S32, which are specifically as follows:

[0064] S31: Inputting the first body shape parameter and the first posture parameter into the parameterized human body model to obtain a human body model.

[0065] In the embodiment of the present application, the length of the corresponding bone in the initial template is scaled according to the first body shape parameter to generate a human skeletal structure that conforms to the specific body shape. The rotation angle of each joint is set according to the first posture parameter.

[0066] S32: Using the forward dynamics method, obtain the second three-dimensional joint point data according to the joint rotation angle of each joint point in the human body model.

[0067] Among them, Forward Kinematics (FK) calculates the position and orientation of the end effector based on known joint angles.

[0068] In an embodiment of the present application, based on the rotation angle of each joint, the three-dimensional position coordinates of each joint are obtained using the forward dynamics method.

[0069] The embodiment of the present application adjusts the parameterized human body model based on the first body shape parameter and the first posture parameter, and can automatically and quickly obtain the second three-dimensional joint point data.

[0070] In one embodiment, the human posture prior distribution model includes an encoder network. The step of obtaining the first latent variable distribution according to the first body shape parameter, the first posture parameter, and the human posture prior distribution model in step S30 includes step S33, which is specifically as follows:

[0071] S33: Compress the first body shape parameter and the first posture parameter into a low-dimensional representation of the latent space through the encoder network to obtain a first latent variable distribution.

[0072] In an embodiment of the present application, the first body shape parameter and the first posture parameter are input into the encoder network, and the encoder network outputs the mean and variance of the latent variable.

[0073] The embodiment of the present application encodes the first body shape parameter and the first posture parameter through an encoder network, and can automatically and quickly obtain the first latent variable distribution.

[0074] In one embodiment, step S40 includes steps S41 to S46, which are specifically as follows:

[0075] S41: Obtain a first loss function value based on the first three-dimensional joint point data, the second three-dimensional joint point data and a preset first loss function.

[0076] Among them, the preset first loss function can be set according to actual needs.

[0077] In an embodiment of the present application, the error between the second three-dimensional joint point data and the first three-dimensional joint point data of each joint point is calculated, the L2 norm of the error is calculated, and the first loss function value is obtained.

[0078] S42: Obtain a second loss function value according to the first latent variable distribution, a preset second latent variable distribution, and a preset second loss function.

[0079] Among them, the preset second loss function can be set according to actual needs.

[0080] In an embodiment of the present application, the mean error and variance error of the first latent variable distribution and the preset second latent variable distribution are calculated, and the L2 norm of the mean error and the variance error is calculated to obtain the second loss function value.

[0081] S43: Obtain a third loss function value according to the first body shape parameter and a preset third loss function.

[0082] Among them, the preset third loss function can be set according to actual needs.

[0083] In the embodiment of the present application, the L2 norm of the first body shape parameter is calculated to obtain the third loss function value.

[0084] S44: Obtain a fourth loss function value according to the second three-dimensional joint point data of each video frame and a preset fourth loss function.

[0085] Among them, the preset fourth loss function can be set according to actual needs.

[0086] In the embodiment of the present application, the error of the second three-dimensional joint point data of the same joint point in any two adjacent video frames is calculated to obtain a plurality of errors. The plurality of errors are summed, and the L2 norm of the summed result is calculated to obtain a fourth loss function value.

[0087] S45: Obtain a total loss function value based on the first loss function value, the second loss function value, the third loss function value, and the fourth loss function value.

[0088] In an embodiment of the present application, a weighted sum is performed on the first loss function value, the second loss function value, the third loss function value, and the fourth loss function value to obtain a total loss function value.

[0089] S46: Update the first body shape parameter and the first posture parameter according to the total loss function value until a preset condition is met, thereby obtaining the second body shape parameter and the second posture parameter.

[0090] In the embodiment of the present application, the expression of the total loss function value is as follows:

[0091]

[0092] Among them, J represents the first three-dimensional joint point data of each joint point, J estRepresents the second three-dimensional joint point data of each joint point, (μ est ,σ est ) represents the mean and variance of the first latent variable distribution, (μ, σ) represents the mean and variance of the preset second latent variable distribution, α shape represents the first body shape parameter, J est,i Represents the second 3D joint point data of each joint point in the i-th video frame, J est,i-1 represents the second 3D joint point data of each joint point in the i-1th video frame, n represents the number of video frames, λ k ,λ piror ,λ shape ,λ vel Represents weight.

[0093] Based on the calculation of the above-mentioned loss function value, the embodiment of the present application utilizes a gradient descent algorithm to iteratively update the first body shape parameter and the first posture parameter to obtain the second body shape parameter and the second posture parameter. Because the loss function takes into account the three-dimensional joint position errors, the human motion prior errors, the small range of variation in the body shape parameters, and the reduction of jitter between frames, the accuracy of the second body shape parameter and the second posture parameter can be improved.

[0094] In one embodiment, step S50 includes steps S51 to S52, which are specifically as follows:

[0095] S51: Obtain a third posture parameter according to the video frame data, the second posture parameter, and the posture jitter smoothing model.

[0096] Among them, the posture jitter smoothing model is trained on a large number of human motion data sets. The training set contains video frame data and corresponding human parametric model data, which can effectively reduce the jitter problem of motion data and achieve a natural transition of motion.

[0097] In an embodiment of the present application, the video frame data and the second posture parameter are input into a posture jitter smoothing model, and the posture jitter smoothing model outputs a third posture parameter after jitter smoothing.

[0098] S52: Obtain optimized three-dimensional joint point data according to the second body shape parameter, the third posture parameter, and the parameterized human body model.

[0099] In an embodiment of the present application, the parameterized human body model is adjusted according to the second body shape parameter and the third posture parameter to obtain optimized three-dimensional joint point data.

[0100] The embodiment of the present application utilizes a posture jitter smoothing model to perform jitter smoothing processing on the second posture parameter, which can improve the quality of the optimized three-dimensional joint point data and avoid jitter in the reconstructed human body movement.

[0101] In one embodiment, the posture jitter smoothing model includes a feature extraction network and a posture smoothing network. Step S51 includes steps S511 to S512, which are specifically as follows:

[0102] S511: Perform feature extraction on the video frame data through a feature extraction network to obtain feature data.

[0103] In an embodiment of the present application, the feature extraction network is a convolutional neural network, and each video frame is input into the convolutional neural network for feature extraction to obtain feature data.

[0104] S512: Input the feature data and the second posture parameter into a posture smoothing network to obtain a third posture parameter.

[0105] In an embodiment of the present application, the posture smoothing network is a Transformer model. The Transformer model is a deep learning model. The Transformer model includes a self-attention mechanism and position encoding. It can process sequence data in parallel, significantly improving training efficiency.

[0106] Specifically, the feature data and the second posture parameter are input into a posture smoothing network, and the second posture parameter is smoothed by the posture smoothing network to obtain a smoothed third posture parameter.

[0107] The embodiment of the present application is based on the cooperation of the feature extraction network and the posture smoothing network, and can automatically and quickly obtain the third posture parameter.

[0108] In one embodiment, after step S50, steps S61 to S66 are included, specifically as follows:

[0109] S61: Acquire three-dimensional joint point data of the foot joint points from the optimized three-dimensional joint point data.

[0110] The foot joints include the left toe joint, the left heel joint, the right toe joint, and the right heel joint.

[0111] In the embodiment of the present application, after obtaining the optimized three-dimensional position coordinates of each joint point, the three-dimensional position coordinates of the foot joint point can be directly obtained.

[0112] S62: Obtain the speed, acceleration, and ground distance of the foot joints based on the three-dimensional joint data of the foot joints.

[0113] In an embodiment of the present application, except for the first video frame (the foot joint is grounded by default in the first video frame and is not calculated), the difference in the three-dimensional position of the foot joint between each video frame and the previous video frame is used as the speed of the foot joint.

[0114] Starting from the third video frame (including the third frame), the difference between the velocity of the foot joint in each video frame and the velocity of the foot joint in the previous video frame is used as the acceleration of the foot joint.

[0115] The difference between the y-axis coordinate value of the foot joint point and the horizontal ground (value 0) is used as the distance of the foot joint point from the ground.

[0116] S63: Obtain a first touchdown probability of the foot joint according to the speed, acceleration, and ground distance of the foot joint.

[0117] In this embodiment of the present application, when the speed and acceleration of a foot joint are too large, or the distance from the ground is too high, the foot joint is considered to be off-ground. Specifically, when the speed of the foot joint is greater than a preset speed threshold, or the acceleration of the foot joint is greater than a preset acceleration threshold, or the distance from the ground is greater than a preset distance threshold, the first probability of the foot joint touching the ground is 0. Otherwise, the first probability of the foot joint touching the ground is 1.

[0118] S64: Obtain a second ground contact probability of the foot joint point based on the feature data, the optimized three-dimensional joint point data, and the foot ground contact prediction model.

[0119] Among them, the foot grounding prediction model is used to predict the grounding probability of the foot joints.

[0120] In an embodiment of the present application, the feature data and the optimized three-dimensional joint point data are input into a foot grounding prediction model, and the foot grounding prediction model outputs the grounding probability of the foot joint point, thereby obtaining a second grounding probability of the foot joint point.

[0121] S65: Multiply the first ground contact probability by the second ground contact probability to obtain a target ground contact probability of the foot joint.

[0122] In the embodiment of the present application, the product of the first touchdown probability and the second touchdown probability of the foot joint is used as the target touchdown probability of the foot joint.

[0123] S66: Obtain optimized three-dimensional joint point data of the foot joints according to the target ground contact probability and the three-dimensional joint point data of the foot joints.

[0124] In an embodiment of the present application, after obtaining the target ground contact probability, it is possible to determine which foot joints in which video frames are grounded, thereby processing the three-dimensional positions of the foot joints in the corresponding video frames to obtain optimized three-dimensional joint data of the foot joints, thereby reducing the sliding phenomenon.

[0125] The embodiment of the present application combines the three-dimensional joint point data of the foot joint points with the foot ground contact prediction model to improve the accuracy of the foot ground contact probability prediction.

[0126] In one embodiment, the foot joints include a left toe joint, a left heel joint, a right toe joint, and a right heel joint. Step S66 includes steps S661 to S663, which are specifically as follows:

[0127] S661: When the target ground contact probability of at least one joint point among the foot joint points is a first preset value, obtain the ground distance of the foot joint point in each video frame; compare the sizes of the ground distances to determine the minimum ground distance; subtract the ground distance of each joint point in each video frame from the minimum ground distance to obtain the updated three-dimensional joint point data of each joint point in each video frame.

[0128] In an embodiment of the present application, the video frames are processed frame by frame. If any of the left toe joint, left heel joint, right toe joint, and right heel joint in the current video frame is grounded, the distance between the y-axis coordinates of these four joints and the ground is calculated, and the minimum value of the distance from the ground is determined. The y-axis coordinates of all the joints in the current video frame are subtracted from the minimum value to obtain the updated three-dimensional joint point data of each joint in the current frame.

[0129] S662: Determine a touchdown foot joint point of each video frame according to the target touchdown probability; wherein the touchdown foot joint point is a foot joint point whose target touchdown probability is a first preset value.

[0130] The first preset value can be set according to actual needs. Specifically, the first preset value is 1.

[0131] In the embodiment of the present application, if the target ground contact probability of the foot joint is 1, the foot joint is a ground contact foot joint.

[0132] S663: When the grounded foot joint point in a continuous preset number of video frames is the same foot joint point, the average value of the updated three-dimensional joint point data of the grounded foot joint point in the continuous preset number of video frames is calculated, and the average value is used as the optimized three-dimensional joint point data of the grounded foot joint point of each video frame in the continuous preset number of video frames.

[0133] In this embodiment of the present application, each foot joint is processed individually. Using the target touchdown probability, it is determined whether the four foot joints (left toe, left heel, right toe, and right heel) are grounded in each frame. Each grounded foot joint is processed separately. Specifically, video frames with consecutive grounding (a target touchdown probability of 1 for more than two frames) are merged into a grounding interval. The three-dimensional coordinates of the grounded foot joints within the interval are averaged across all dimensions (x, y, and z axes). The original data within the grounding interval is replaced with the average value, resulting in a unified average value for the three-dimensional coordinates of the grounded foot joints within the interval, thus obtaining optimized three-dimensional joint data for the grounded foot joints.

[0134] The embodiment of the present application calculates the average value of each dimension of the three-dimensional coordinates of the ground contact foot joint point within the contact area, and replaces the original data within the ground contact area with the average value, thereby reducing the slipping phenomenon.

[0135] In one embodiment, after the step of using the average value as the optimized three-dimensional joint point data of the ground contact foot joint of each video frame in the consecutive preset number of video frames in step S663, the method includes step S6631, which is specifically as follows:

[0136] S6631: Perform window smoothing on the optimized 3D joint point data of the foot joints.

[0137] Among them, the window smoothing algorithm refers to a method of smoothing data by calculating the average value of the data within the window, which can effectively reduce the random noise in the data while retaining the overall trend of the data.

[0138] In an embodiment of the present application, in order to make the foot movement transition smooth and avoid the foot movement transition being too stiff, a window sliding algorithm is used to smoothly adjust the three-dimensional joint point data after the foot joint points are optimized.

[0139] In one embodiment, after the step of using the average value as the optimized three-dimensional joint point data of the grounded foot joint in each of the preset number of consecutive video frames in step S663, the method includes steps S6632 to S6633, which are specifically as follows:

[0140] S6632: Obtain optimized 3D joint point data of the foot joint points of the last video frame among a preset number of consecutive video frames and optimized 3D joint point data of the foot joint points of the video frame next to the last video frame;

[0141] S6633: Perform Gaussian filtering on the optimized three-dimensional joint point data of the foot joint points of the last video frame and the optimized three-dimensional joint point data of the foot joint points of the next video frame after the last video frame.

[0142] Among them, the Gaussian filtering algorithm is a linear smoothing filtering method that uses a Gaussian function to replace each point in the original data and calculates the new data value by weighted averaging.

[0143] In an embodiment of the present application, a Gaussian filtering algorithm is used to perform Gaussian filtering on the last video frame between the grounding areas and the subsequent ungrounded video frames (video frames in the ungrounded interval) to avoid excessive changes in the motion between the grounding interval and the ungrounded video frames, which affects the visual effect.

[0144] The following are embodiments of the apparatus of the present application, which can be used to perform the content of the method of the present application. For details not disclosed in the embodiments of the apparatus of the present application, please refer to the content of the method in the embodiments of the present application.

[0145] See Figure 2 , which shows a schematic diagram of the structure of the motion capture data optimization device provided in an embodiment of the present application. The motion capture data optimization device 6 provided in an embodiment of the present application includes:

[0146] The video frame data acquisition module 61 is used to acquire video frame data; wherein the video frame data is used to record human body movements;

[0147] A joint point data acquisition module 62 is configured to obtain first three-dimensional joint point data based on the video frame data;

[0148] Latent variable distribution acquisition module 63 is used to obtain first body shape parameters and first posture parameters; obtain second three-dimensional joint point data based on the first body shape parameters, first posture parameters, and a parameterized human body model; and obtain a first latent variable distribution based on the first body shape parameters, first posture parameters, and a human body posture prior distribution model;

[0149] a posture parameter obtaining module 64 for iteratively updating the first body shape parameter and the first posture parameter according to the first three-dimensional joint point data, the second three-dimensional joint point data, the first latent variable distribution, and a preset second latent variable distribution until a preset condition is satisfied, thereby obtaining a second body shape parameter and a second posture parameter;

[0150] The joint point data optimization module 65 is used to obtain optimized three-dimensional joint point data according to the second body shape parameter, the second posture parameter and the parameterized human body model.

[0151] The embodiment of the present application is applied, by obtaining video frame data; wherein the video frame data is used to record human body movements; obtaining first three-dimensional joint point data based on the video frame data; obtaining first body shape parameters and first posture parameters; obtaining second three-dimensional joint point data based on the first body shape parameters, the first posture parameters and a parameterized human body model; obtaining a first latent variable distribution based on the first body shape parameters, the first posture parameters and a human body posture prior distribution model; iteratively updating the first body shape parameters and the first posture parameters based on the first three-dimensional joint point data, the second three-dimensional joint point data, the first latent variable distribution and a preset second latent variable distribution until a preset condition is met, thereby obtaining second body shape parameters and second posture parameters; obtaining optimized three-dimensional joint point data based on the second body shape parameters, the second posture parameters and the parameterized human body model. The present application updates the first body shape parameters and the first posture parameters in the iteration so that the first latent variable distribution obtained based on the updated first body shape parameters and the updated first posture parameters conforms to the preset second latent variable distribution, thereby ensuring that the human body movements reconstructed based on the updated first posture parameters are within the normal range of human activities, avoiding the irrationality of motion capture data and improving the quality of motion capture data. Since the reconstructed human body movements are no longer distorted and do not require manual adjustment, the efficiency of movement output is improved and the cost of restoration is reduced.

[0152] The following are device embodiments of the present application, which can be used to perform the content of the method in the embodiment of the present application. For details not disclosed in the device embodiments of the present application, please refer to the content of the method in the embodiment of the present application.

[0153] See also Figure 3 The present application also provides an electronic device 300, which can be specifically a computer, a mobile phone, a tablet computer, a motion capture data optimization device, etc. In an exemplary embodiment of the present application, the electronic device 300 is a motion capture data optimization device, which can include: at least one processor 301, at least one memory 302, at least one display, at least one network interface 303, a user interface 304 and at least one communication bus 305.

[0154] The user interface 304 is mainly used to provide an input interface for the user and obtain data input by the user. Optionally, the user interface can also include a standard wired interface or a wireless interface.

[0155] The network interface 303 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface).

[0156] The communication bus 305 is used to realize the connection and communication between these components.

[0157] Among them, the processor 301 may include one or more processing cores. The processor uses various interfaces and lines to connect the various parts of the entire electronic device, and performs various functions of the electronic device and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory, and calling data stored in the memory. Optionally, the processor can be implemented in the form of at least one hardware of digital signal processing (DSP), field programmable gate array (FPGA), and programmable logic array (PLA). The processor can integrate one or more combinations of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing the content to be displayed by the display layer; and the modem is used to handle wireless communications. It is understandable that the above-mentioned modem may not be integrated into the processor and may be implemented separately through a chip.

[0158] Among them, the memory 302 may include a random access memory (RAM) or a read-only memory (Read-Only Memory). Optionally, the memory includes a non-transitory computer-readable storage medium. The memory can be used to store instructions, programs, codes, code sets or instruction sets. The memory may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data involved in the above-mentioned various method embodiments, etc. The memory may also be optionally at least one storage device located away from the aforementioned processor. As Figure 3 As shown, the memory as a computer storage medium may include an operating system, a network communication module, a user interface module, and an operating application program.

[0159] The processor can be used to call the application of the motion capture data optimization method stored in the memory, and specifically execute the method steps of the above-mentioned embodiment. The specific execution process can be found in the specific description shown in the embodiment, which will not be repeated here.

[0160] This application also provides a computer-readable storage medium having a computer program stored thereon, with instructions suitable for being loaded by a processor and executing the method steps of the above-described embodiments. The specific execution process can be referred to the specific description of the embodiments and is not described in detail here. The device where the storage medium is located can be a personal computer, laptop computer, smartphone, tablet computer, or other electronic device.

[0161] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are merely illustrative, wherein the components described as separate parts may or may not be physically separated, and the parts shown as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present application scheme. A person of ordinary skill in the art can understand and implement it without paying any creative work.

[0162] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0163] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including the instruction device, which implements the function selected in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 function selected in a box or multiple boxes.

[0164] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 steps for the function selected in a box or multiple boxes.

[0165] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0166] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0167] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be used to store information using any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, tape, disk or other magnetic storage device or any other non-transmission medium that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0168] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0169] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.

Claims

1. A motion capture data optimization method, characterized in that: The steps include: Acquire video frame data; wherein the video frame data is used to record human body movements; Obtaining first three-dimensional joint point data according to the video frame data; Obtaining first body shape parameters and first posture parameters; obtaining second three-dimensional joint point data based on the first body shape parameters, the first posture parameters, and a parameterized human body model; obtaining a first latent variable distribution based on the first body shape parameters, the first posture parameters, and a human body posture prior distribution model; Iteratively updating the first body shape parameter and the first posture parameter according to the first three-dimensional joint point data, the second three-dimensional joint point data, the first latent variable distribution, and a preset second latent variable distribution until a preset condition is satisfied, thereby obtaining a second body shape parameter and a second posture parameter; Optimized three-dimensional joint point data is obtained based on the second body shape parameter, the second posture parameter, and the parameterized human body model.

2. The motion capture data optimization method according to claim 1, wherein: The video frame data includes a plurality of video frames of human body movements captured by multiple cameras; The step of obtaining first three-dimensional joint point data according to the video frame data includes: Extracting two-dimensional position information of each joint point from each of the video frames; Obtaining three-dimensional position information of each joint point based on the two-dimensional position information and parameter information of the corresponding camera; The three-dimensional position information of each joint point is subjected to stereo triangulation processing to obtain first three-dimensional joint point data of each joint point.

3. The motion capture data optimization method according to claim 1, wherein: The step of obtaining second three-dimensional joint point data according to the first body shape parameter, the first posture parameter, and the parameterized human body model includes: Inputting the first body shape parameter and the first posture parameter into the parameterized human body model to obtain a human body model; The forward dynamics method is used to obtain second three-dimensional joint point data according to the joint rotation angle of each joint point in the human body model.

4. The motion capture data optimization method according to claim 1, wherein: The human body posture prior distribution model includes an encoder network; The step of obtaining a first latent variable distribution according to the first body shape parameter, the first posture parameter, and a human body posture prior distribution model includes: The first body shape parameter and the first posture parameter are compressed into a low-dimensional representation of a latent space through the encoder network to obtain a first latent variable distribution.

5. The motion capture data optimization method according to claim 1, wherein: The step of updating the first body shape parameter and the first posture parameter according to the first three-dimensional joint point data, the second three-dimensional joint point data, the first latent variable distribution, and a preset second latent variable distribution until a preset condition is satisfied to obtain a second body shape parameter and a second posture parameter includes: Obtaining a first loss function value according to the first three-dimensional joint point data, the second three-dimensional joint point data, and a preset first loss function; Obtaining a second loss function value according to the first latent variable distribution, a preset second latent variable distribution, and a preset second loss function; Obtaining a third loss function value according to the first body shape parameter and a preset third loss function; Obtaining a fourth loss function value according to the second three-dimensional joint point data of each video frame and a preset fourth loss function; Obtaining a total loss function value according to the first loss function value, the second loss function value, the third loss function value, and the fourth loss function value; According to the total loss function value, the first body shape parameter and the first posture parameter are updated until a preset condition is met, thereby obtaining a second body shape parameter and a second posture parameter.

6. The motion capture data optimization method according to claim 1, wherein: The step of obtaining optimized three-dimensional joint point data according to the second body shape parameter, the second posture parameter, and the parameterized human body model comprises: Obtaining a third posture parameter according to the video frame data, the second posture parameter, and a posture jitter smoothing model; Optimized three-dimensional joint point data is obtained based on the second body shape parameter, the third posture parameter, and the parameterized human body model.

7. The motion capture data optimization method according to claim 6, wherein: The posture jitter smoothing model includes a feature extraction network and a posture smoothing network; The step of obtaining a third posture parameter according to the video frame data, the second posture parameter, and the posture jitter smoothing model includes: Performing feature extraction on the video frame data through the feature extraction network to obtain feature data; The feature data and the second posture parameter are input into the posture smoothing network to obtain a third posture parameter.

8. The motion capture data optimization method according to claim 7, wherein: After the step of obtaining optimized three-dimensional joint point data according to the second body shape parameter, the second posture parameter and the parameterized human body model, the method includes: Acquiring three-dimensional joint point data of foot joint points from the optimized three-dimensional joint point data; Obtaining a velocity, an acceleration, and a ground distance of the foot joints based on the three-dimensional joint data of the foot joints; Obtaining a first touchdown probability of the foot joint according to the speed, acceleration, and ground distance of the foot joint; Obtaining a second ground contact probability of the foot joint point based on the feature data, the optimized three-dimensional joint point data, and a foot ground contact prediction model; multiplying the first ground contact probability by the second ground contact probability to obtain a target ground contact probability of the foot joint; According to the target ground contact probability and the three-dimensional joint point data of the foot joint point, the optimized three-dimensional joint point data of the foot joint point is obtained.

9. The motion capture data optimization method according to claim 8, wherein: The foot joints include a left toe joint, a left heel joint, a right toe joint, and a right heel joint; The step of obtaining optimized three-dimensional joint point data of the foot joint point according to the target ground contact probability and the three-dimensional joint point data of the foot joint point comprises: When the target ground contact probability of at least one of the foot joints is a first preset value, obtaining the ground distance of the foot joint in each video frame; comparing the ground distances to determine a minimum ground distance; and subtracting the ground distance of each joint in each video frame from the minimum ground distance to obtain updated three-dimensional joint point data for each joint in each video frame; Determining a touchdown foot joint point of each video frame according to the target touchdown probability; wherein the touchdown foot joint point is a foot joint point for which the target touchdown probability is the first preset value; When the ground contact foot joint point in a continuous preset number of video frames is the same foot joint point, the average value of the updated three-dimensional joint point data of the ground contact foot joint point in the continuous preset number of video frames is calculated, and the average value is used as the optimized three-dimensional joint point data of the ground contact foot joint point in each of the continuous preset number of video frames.

10. The motion capture data optimization method according to claim 9, wherein: After the step of using the average value as the optimized three-dimensional joint point data of the ground contact foot joint of each of the video frames in the continuous preset number of video frames, the method includes: Window smoothing is performed on the optimized three-dimensional joint point data of the foot joint points.

11. The motion capture data optimization method according to claim 9, wherein: After the step of using the average value as the optimized three-dimensional joint point data of the ground contact foot joint of each of the video frames in the continuous preset number of video frames, the method includes: Obtaining optimized three-dimensional joint point data of the foot joint points of the last video frame of the preset number of consecutive video frames and optimized three-dimensional joint point data of the foot joint points of the video frame next to the last video frame; Gaussian filtering is performed on the optimized three-dimensional joint point data of the foot joint points of the last video frame and the optimized three-dimensional joint point data of the foot joint points of the next video frame after the last video frame.

12. A motion capture data optimization device, characterized in that: include: A video frame data acquisition module, configured to acquire video frame data, wherein the video frame data is used to record human body movements; A joint point data acquisition module, configured to obtain first three-dimensional joint point data based on the video frame data; a latent variable distribution acquisition module, configured to acquire a first body shape parameter and a first posture parameter; obtain second three-dimensional joint point data based on the first body shape parameter, the first posture parameter, and a parameterized human body model; and obtain a first latent variable distribution based on the first body shape parameter, the first posture parameter, and a human body posture prior distribution model; a posture parameter acquisition module, configured to iteratively update the first body shape parameter and the first posture parameter according to the first three-dimensional joint point data, the second three-dimensional joint point data, the first latent variable distribution, and a preset second latent variable distribution until a preset condition is satisfied, thereby obtaining a second body shape parameter and a second posture parameter; The joint point data optimization module is used to obtain optimized three-dimensional joint point data based on the second body shape parameter, the second posture parameter and the parameterized human body model.

13. An electronic device, characterized in that: include: A processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the motion capture data optimization method according to any one of claims 1 to 11.

14. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the motion capture data optimization method according to any one of claims 1 to 11 is implemented.